Skip to content

Repository files navigation

GenTest

A competitive programming testcase generator with ordinary Python configs, reproducible inputs, and checked Python/C++ reference solutions. You control every input format and generation rule.

Quick Start

Install Python 3.10 or newer. C++ references also need g++ on your PATH. The core CLI has no third-party runtime dependencies.

git clone https://github.com/khanhtran0111/GenTest.git
cd GenTest
python -m gentest new sumarray --template array --lang python

Open problems/sumarray/config.py in your editor. Adjust the test count, constraints and strategies. Open problems/sumarray/reference.py and paste your correct solution, reading standard input and writing standard output.

The array template already includes a working example reference: read n and an array, then print its sum. You can use it immediately to learn the workflow:

python -m gentest validate sumarray --seed 2026
python -m gentest generate sumarray --seed 2026 --zip

Inspect problems/sumarray/tests/Test001/sumarray.inp and sumarray.out. The first input is 1 followed by 0; its output is 0. The manifest is problems/sumarray/tests/manifest.json, and the archive is problems/sumarray/tests.zip.

To regenerate the set:

python -m gentest generate sumarray --seed 2026 --clean --zip

On systems where Python is named python3, use python3 in these commands. On Windows, py also works. GenTest executes Python references using the same interpreter you used to launch it.

For guided setup, run python -m gentest with no arguments. It asks for the problem name, template, language, test count and number of subtasks. It creates the files and prints the next command; edit your constraints before generation.

flowchart LR
    accTitle: First Test Set Workflow
    accDescr: Create a problem, edit its Python config, supply a reference, generate and inspect tests.
    create[Create problem] --> config[Edit config.py]
    config --> reference[Paste reference solution]
    reference --> generate[Generate with a seed]
    generate --> inspect[Inspect inputs and outputs]
Loading

Installation and commands

Working directly from the repository needs no installation. Optionally install the CLI into your current Python environment:

python -m pip install -e .
gentest doctor
gentest list-templates

gentest and python -m gentest expose the same commands. Run them from the project root, or put --root /path/to/project before the command.

Command Purpose
new <problem> --template array --lang python Create a config and exactly one reference file; never overwrite an existing problem
generate <problem> Generate and validate inputs, execute references, publish a complete test set
validate <problem> Check config, reference settings/path, generated inputs, validators and duplicates; no compilation, execution or output files
list-templates List available starter configs
doctor Show Python, project root and whether g++ is available
Generation option Meaning
--seed 2026 Global integer seed; omitted seeds are generated and printed
--clean Replace the previous tests/ after the whole run succeeds
--zip Write tests.zip containing TestXXX/ directories and the manifest
--tests 30 Override the total count; explicit strategy/subtask counts must still fit
--verbose Show each test's metadata and full framework tracebacks on errors

validate also accepts --seed, --tests and --verbose. new accepts --tests to set the initial count. Problem names use letters, digits, underscores and hyphens and must be valid on Windows.

Existing output requires --clean. A failed run stops at the first failed testcase, publishes no new test set, and leaves a previous successful set intact. That previous set's manifest still describes its original seed. Successful regeneration without --zip removes an older tests.zip to avoid a stale archive. Compilation uses a temporary build directory, and each output file is written atomically only after successful reference execution and any cross-check.

Project structure

GenTest/
  gentest/
    cli.py              # Commands and beginner wizard
    config.py           # Python config loading and distribution checks
    engine.py           # Seeds, generation, validation, manifests and exports
    runner.py           # Checked Python/C++ execution
    utils.py            # Generation and formatting helpers
    templates/          # Editable Python starter configs
  problems/
    sumarray/
      config.py
      reference.py      # Or reference.cpp, selected explicitly in config
      brute.py          # Optional second reference
      tests/
        manifest.json
        Test001/
          sumarray.inp
          sumarray.out
      tests.zip         # Optional export
  solutions/            # Optional legacy <problem>.py or <problem>.cpp
  config_sample.py      # Original legacy example
  config_sample/        # Original legacy templates, including corrected graphs
  genTest.py            # Legacy batch generation launcher
  genUltils.py          # Compatibility import of gentest.utils
  app.py                # Alternate launcher for the same CLI/wizard
  tests/                # Automated tests for GenTest itself

new creates problems/ as needed. A fresh checkout does not need a solutions/ directory. Generated test sets, archives, bytecode and build artifacts are ignored by Git; configs and reference source files can be committed.

Your config is Python

Use any functions, imports, loops or custom data structures needed by the problem. The engine does not interpret problem statements. Configs execute as Python code in your environment.

The modern generator returns the entire input as a string:

def generate(ctx):
    n = ctx.rng.randint(1, 100)
    return f"{n}\n"
Context field Meaning
ctx.test_id 1-based test ID, used in Test001 etc.
ctx.subtask 1-based subtask index
ctx.strategy One of the six strategy names below
ctx.seed Deterministic per-test seed
ctx.rng An independent random.Random for this testcase
ctx.strategy_index 1-based occurrence of this strategy across the run

problemName determines .inp/.out filenames. totalOfTests determines how many test directories are created. solution selects the reference explicitly. Relative solution paths are resolved from the directory containing config.py, including paths with spaces.

Subtasks

subtasks = [30, 70] allocates 30% of the test count to subtask 1 and 70% to subtask 2. These are test allocation percentages; GenTest does not assign judge scores. Percentages must be finite, nonnegative and sum to 100.

Allocation uses cumulative floors: with 7 tests and [50, 50], subtask 1 gets 3 tests and subtask 2 gets 4. A small percentage can receive no tests. Zero-percent subtasks are skipped correctly. Use enough tests to cover each intended subtask.

Choose constraints inside the generator, for example maxN = [100, 100000] and limit = maxN[ctx.subtask - 1]. The wizard initially assigns nearly equal percentages and identical limits; customize both.

For exact counts, add subtask_counts = [3, 7] with totalOfTests = 10. Keep one count for each entry in subtasks; counts must sum to the total. This overrides percentage rounding while the percentage list must still be valid.

Six testcase strategies

Strategy What you control
fixed Hand-written regression or sample inputs in fixed_cases, or custom fixed logic
boundary Minimum/maximum sizes and values, empty structures when permitted, endpoints
special All equal, sorted, periodic, disconnected, star-shaped or other meaningful structures
random Broad randomized coverage with problem-specific distributions
adversarial Inputs aimed at likely wrong assumptions or poor complexity
stress Maximum constraints for time and memory pressure

Strategies are labels and dispatch controls, not automatic test-quality guarantees. Your generator decides what each means.

fixed_cases = ["1\n0\n", "2\n-1 1\n"]
strategies = {
    "fixed": 2,
    "boundary": 2,
    "special": 2,
    "adversarial": 2,
    "random": "remaining",
    "stress": 2,
}

Groups run in dictionary order. Only one group may use "remaining"; it fills whatever is left after all numeric counts. Fixed inputs bypass generate, but still pass through validation and duplicate checking. If fixed_cases is nonempty, its length must equal the fixed count. With no explicit strategy settings, fixed inputs come first and random fills the rest.

The subtask and strategy sequences are allocated independently. To give each subtask its own deliberate strategy mix, use test_plan instead of strategies:

totalOfTests = 10
subtasks = [50, 50]
test_plan = [
    {"subtask": 1, "strategy": "boundary", "count": 1},
    {"subtask": 1, "strategy": "random", "count": 4},
    {"subtask": 2, "strategy": "adversarial", "count": 2},
    {"subtask": 2, "strategy": "stress", "count": 3},
]

Plan counts must match both the total and each subtask's allocation. Remove unused fixed_cases when replacing a template's strategy plan. If an override such as --tests 3 cannot fit the configured groups, GenTest reports the inconsistency; edit the config or create a smaller starter with new --tests 3.

Seed and reproducibility

Every run prints its global seed before loading the config. Each test seed is the first 8 bytes, interpreted as an unsigned big-endian integer, of SHA-256 over the ASCII string global_seed:test_id.

The manifest records the problem, generator version, global seed, number of tests, and each test's ID, subtask, strategy, per-test seed and SHA-256 input hash. Inputs are UTF-8 bytes, written without platform newline conversion.

Keep the config, GenTest/Python versions and seed together for reproducibility. Use ctx.rng for all randomness; helpers accept rng=ctx.rng. Avoid time, OS randomness, unordered iteration, and mutable state shared between testcases. Imported third-party RNGs and external data sources remain under your control. The engine seeds standard-library global random separately for each testcase to support legacy configs, then restores its previous state. It also seeds global random while loading the config.

Regenerate a whole set with the manifest's global seed, not a reported per-test seed:

python -m gentest generate sumarray --seed 2026 --clean

For an individual testcase, use its manifest entry and rebuild the context. Load the config with the original global seed so any import-time standard-library randomness also matches:

import json
import random
from pathlib import Path
from gentest.config import load_config
from gentest.engine import Context, input_content

folder = Path("problems/sumarray")
manifest = json.loads((folder / "tests/manifest.json").read_text())
item = manifest["tests"][6]  # Test007
config = load_config(folder / "config.py", seed=manifest["global_seed"])
index = sum(t["strategy"] == item["strategy"] for t in manifest["tests"][:7])
ctx = Context(item["test_id"], item["subtask"], item["strategy"],
              item["seed"], random.Random(item["seed"]), index)
print(input_content(config, ctx), end="")

The same construction works for a failed testcase using its printed metadata and the strategy occurrence from your plan. Generation is sequential; configs should not start threads that depend on global random.

Built-in templates

Each starter contains comments on constraints, boundaries, random coverage, special/adversarial structures and stress cases. Edit them for your actual problem. All templates except array create an intentionally failing reference placeholder, so an unfinished solution cannot silently produce empty answer files.

CLI template Input and design examples
single-number One integer; minimum, maximum, powers of two, near-maximum values
array n and integers; singleton, equal values, alternating extremes, maximum length
string Lowercase string; singleton, repeated/periodic patterns, long common prefix
multiple-test-cases T arrays; tiny cases, repeated cases, mixed sizes, maximum total size
matrix Integer matrix; single cell, narrow matrix, zero and alternating matrices
grid ./# grid; open, blocked, checkerboard, maximum area
permutation Permutation of 1..n; identity, reverse, displaced minimum, shuffled
intervals Inclusive ranges; points, full spans, nested and random overlaps
queries Point updates/range queries; endpoint indices, repeated updates, large workloads
tree Undirected tree; singleton, path, star, random parents
undirected-graph Simple graph; no edges, paths, random density, complete graph
directed-graph Simple directed graph; paths, a cycle, random edges, all ordered pairs
weighted-graph Simple weighted graph; endpoint uniqueness, low/high weights, dense stress
dag Acyclic directed graph; chains, shuffled topological order, complete DAG

Complete array example

The task is to sum an array. Save this as problems/sumarray/config.py after creating the problem:

from gentest import utils

problemName = "sumarray"
totalOfTests = 12
subtasks = [50, 50]
maxN = [20, 10000]
solution = {"path": "reference.py", "language": "python", "timeout": 2.0}
fixed_cases = ["1\n0\n", "2\n-1000000 1000000\n"]
strategies = {"fixed": 2, "boundary": 2, "special": 1,
              "adversarial": 1, "random": "remaining", "stress": 2}
duplicate_policy = "warn"


def generate(ctx):
    limit = maxN[ctx.subtask - 1]
    n = limit if ctx.strategy in ("boundary", "stress") else ctx.rng.randint(1, limit)
    if ctx.strategy == "boundary":
        a = [10**6 if ctx.strategy_index % 2 else -10**6] * n
    elif ctx.strategy == "special":
        a = [0] * n
    elif ctx.strategy == "adversarial":
        a = [(-1)**i * 10**6 for i in range(n)]
    else:
        a = utils.integers(n, -10**6, 10**6, rng=ctx.rng)
    return utils.toStringLines([n, a]) + "\n"


def validateInput(content, ctx):
    n, *a = map(int, content.split())
    if len(a) != n or not 1 <= n <= maxN[ctx.subtask - 1]:
        raise ValueError("Wrong array length for this subtask")
    if any(abs(x) > 10**6 for x in a):
        raise ValueError("Array element out of range")

The complete reference.py:

import sys
n, *a = map(int, sys.stdin.buffer.read().split())
print(sum(a))

Run python -m gentest generate sumarray --seed 2026 --clean. Large values can make the sum exceed a 32-bit integer; a C++ reference should use long long.

Complete graph example

This example asks for the number of connected components in a simple undirected graph. Create the files with:

python -m gentest new components --template undirected-graph --lang python

Replace problems/components/config.py with:

from gentest import utils

problemName = "components"
totalOfTests = 10
subtasks = [100]
solution = {"path": "reference.py", "language": "python", "timeout": 2.0}
fixed_cases = ["1 0\n", "4 2\n1 2\n3 4\n"]
strategies = {"fixed": 2, "boundary": 1, "special": 1,
              "adversarial": 1, "random": "remaining", "stress": 1}


def generate(ctx):
    n = 100 if ctx.strategy == "stress" else ctx.rng.randint(2, 30)
    if ctx.strategy == "boundary":
        edges = []
    elif ctx.strategy == "special":
        edges = utils.tree(n, shape="star", rng=ctx.rng)
    elif ctx.strategy == "adversarial":
        edges = [(u, u + 1) for u in range(1, n - 1)]  # One isolated vertex.
    else:
        m = n * (n - 1) // 2 if ctx.strategy == "stress" else ctx.rng.randint(0, n * (n - 1) // 2)
        edges = utils.graph(n, m, rng=ctx.rng)
    return utils.toStringLines([f"{n} {len(edges)}", *edges]) + "\n"


def validateInput(content, ctx):
    rows = [list(map(int, row.split())) for row in content.splitlines()]
    n, m = rows[0]
    edges = rows[1:]
    if not 1 <= n <= 100 or len(edges) != m:
        raise ValueError("Invalid graph size")
    pairs = set()
    for u, v in edges:
        if not 1 <= u <= n or not 1 <= v <= n or u == v:
            raise ValueError("Invalid endpoint or self-loop")
        pairs.add(tuple(sorted((u, v))))
    if len(pairs) != m:
        raise ValueError("Parallel edges are not allowed")

The complete reference.py uses a disjoint-set structure:

import sys
it = iter(map(int, sys.stdin.buffer.read().split()))
n, m = next(it), next(it)
parent = list(range(n + 1))


def find(x):
    while parent[x] != x:
        parent[x] = parent[parent[x]]
        x = parent[x]
    return x


for _ in range(m):
    u, v = next(it), next(it)
    parent[find(u)] = find(v)
print(len({find(v) for v in range(1, n + 1)}))

Run python -m gentest generate components --seed 42. The fixed cases' outputs are 1 and 2.

Utility reference

Import from gentest import utils and pass rng=ctx.rng to random helpers. Bounds are inclusive, sizes are nonnegative, and graph vertices are numbered 1..n.

Helpers Purpose
integer(low, high), integers(n, low, high) Integer or array; integers supports distinct=True, sort="asc"/"desc"
permutation(n, start=1) Shuffled consecutive integers
distinct_array, sorted_array(..., reverse=True) Unique or sorted/reversed data
duplicate_array(n, low, high, unique=3) Many duplicates from a small sampled value pool
string(n, alphabet="abc"), binary_string(n) Custom-alphabet or binary strings
matrix(rows, cols, low, high), grid(rows, cols, alphabet=".#") Integer matrices or character rows
intervals(n, low, high) Inclusive (left, right) pairs with left <= right
queries(n, factories) Select a callable per query; each factory takes an RNG and returns your row
tree(n, shape="random") Random-parent, path or star tree
forest(n, components=3) Acyclic forest with exactly the requested component count
graph(n, m) Configurable graph helper
undirected_graph, directed_graph, connected_graph, weighted_graph Convenience graph variants
dag(n, m, connected=False) DAG oriented along a shuffled topological order
toStringLines, toStringSpace, toStringLinesAll Row formatting, space formatting, or recursive line flattening

graph options are connected, directed, weighted, allow_self_loops, allow_parallel_edges, weight_range=(1, 100), and rng. All boolean options default to false. Directed connected=True means weakly connected, not strongly connected. Trees/forests/DAGs also accept weighted and weight_range.

Simple graphs sample edge ranks without replacement, using O(n + m) memory rather than constructing all O(n²) potential pairs. Weighted edge uniqueness depends on endpoints, not weights. Connected generation first creates a spanning tree. These helpers do not promise a uniform distribution over all connected graphs or labelled trees. Impossible edge counts fail immediately.

Validators and brute-force cross-checking

An optional validateInput(content, ctx) runs before either reference. Return None/True for success, return False to reject, or raise ValueError with a specific explanation. It also runs for fixed inputs and during validate.

Set duplicate_policy = "warn" (default), "error", or "allow". Duplicate detection compares SHA-256 input hashes across the whole run. Warnings identify the current and earlier test IDs; failures include strategy, subtask and seed.

For C++ configuration:

solution = {
    "path": "reference.cpp",
    "language": "cpp",
    "timeout": 2.0,
    "cpp_standard": "c++17",
    "compile_flags": ["-O2", "-Wall"],
    "compiler": "g++",  # Optional compiler command/path.
}

C++ compiles once per run, with a 60-second compilation timeout. Each reference execution has its own timeout. Top-level timeout, cpp_standard and compile_flags provide defaults; solution dictionary entries override them. References execute with their source directory as the working directory and communicate through stdin/stdout. Outputs are captured as bytes; a crash or timeout never becomes an answer file.

To compare an optimized reference with a brute-force solution, add:

brute_solution = {"path": "brute.py", "language": "python", "timeout": 5.0}
cross_check_strategies = ["fixed", "boundary", "random"]


def cross_check(ctx):
    return ctx.subtask == 1  # Keep the brute-force workload small.

The default cross-check strategies are fixed, boundary and random. cross_check(ctx) is an optional additional filter; configure small constraints for selected tests. Comparison uses whitespace-separated output tokens, so spaces and line breaks may differ, but numeric formatting must agree. It is not a floating-point checker. The first mismatch reports the test ID, strategy, subtask, seed, input, and both outputs, and stops generation. Arbitrary problem-specific checks can also live in your own Python code.

How to design a strong test set

Start with small hand-written cases whose answers you understand. Cover every legal boundary: one element, zero edges, extreme values, endpoint indices and the empty case if your statement allows it.

Add special structures that exercise different paths through a solution: equal values, duplicates, increasing/decreasing arrays, periodic strings, disconnected graphs, stars and chains. Add adversarial cases aimed at common mistakes: integer overflow, off-by-one indices, state retained between cases, assumptions about connectivity, deep recursion, or an algorithm that becomes quadratic on a particular ordering.

Use random cases to explore combinations you would not write manually. Vary sizes, value ranges and graph densities rather than choosing every parameter uniformly over its largest range. Check small random cases against a brute-force implementation.

Finish with maximum-constraint stress cases for time and memory. A useful suite combines all of these categories; a large pile of random inputs alone does not establish that a solution is correct. Review the manifest to confirm the intended strategy and subtask distribution actually ran.

Legacy compatibility

Existing configs continue to support:

import random
import genUltils
problemName = "oldproblem"
totalOfTests = 20
subtasks = [50, 50]


def genInputContent(testID, curSubtask):
    return str(random.randint(1, 100))

testID and curSubtask remain 1-based. The old helper names and call signatures remain in genUltils; new code can use gentest.utils. Date generation now handles December correctly, and distinct rounded floats/dates use bounded sampling instead of potentially endless retry loops. Graph samples now print one edge per line and reject repeated endpoint pairs regardless of weight.

For legacy configs without solution, GenTest looks in solutions/<problemName>.py and .cpp. Exactly one must exist. If both exist, explicitly select solution, or set solution_language = "cpp"/"python". No silent Python preference remains. If both generator functions exist, generate(ctx) takes precedence over genInputContent.

python genTest.py generates all repository problems/*/config.py in sorted order, replacing successful prior sets and exporting ZIPs as the old batch command did. It stops at the first failure. With arguments it forwards to the canonical CLI, for example python genTest.py generate oldproblem --seed 42 --clean.

python app.py opens the same beginner wizard; it no longer bulk-creates both languages from name.txt. The original empty name.txt is retained as a legacy file. Old internal GUI/app classes and unchecked generation functions are not retained as public APIs. There is no GUI dependency. Output moves to problems/<folder>/tests/TestXXX/, eliminating the duplicated problem-name directory. The .inp and .out filenames retain the original style.

Troubleshooting

Problem What to check
Missing g++ Run python -m gentest doctor. Install a C++ compiler providing g++ and add its executable directory to PATH, or select Python.
C++ compile error Read the captured compiler stderr. Check syntax, included headers, cpp_standard and compile_flags.
Reference solution TLE The error shows the test and timeout. Check for infinite loops and excessive complexity, or set an appropriate solution['timeout'].
Runtime error Read the exit status and stderr. Reproduce using the reported test/global seed; check input parsing, invalid indexing and recursion depth.
Invalid config Check Python syntax, positive total count, percentages summing to 100, and strategy/subtask counts matching the total.
Validator rejection Read the validator's explanation with the test metadata. Correct the generator or the validator's interpretation of the statement.
Both Python and C++ files exist Select the intended path/language in solution instead of relying on discovery.
Test directory already exists Add --clean to replace the generated set after a successful run.
Module not found Run from the repository root, or install with python -m pip install -e . in the interpreter you use.
Need a traceback Repeat generate or validate with --verbose.

Development

Run the dependency-free test suite:

python -m unittest discover -s tests -v

Tests cover seeds/hashes, independent reproduction, subtask and strategy validation, duplicate policies, validators, transactional replacement, Python/C++ execution, failures/timeouts, cross-checking, graph invariants, CLI/wizard/templates and legacy configs. C++ execution tests skip if g++ is unavailable. CI runs Python 3.10–3.14 on Ubuntu and Windows.

Original authors: Khanh Tran and Ngat Do, Code Dream Programming Learning Center. Licensed under the MIT license.

About

Generate testcase for CP

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages