A competitive programming testcase generator with ordinary Python configs, reproducible inputs, and checked Python/C++ reference solutions. You control every input format and generation rule.
Install Python 3.10 or newer. C++ references also need g++ on your PATH. The core CLI has no third-party runtime dependencies.
git clone https://github.com/khanhtran0111/GenTest.git
cd GenTest
python -m gentest new sumarray --template array --lang pythonOpen problems/sumarray/config.py in your editor. Adjust the test count, constraints and strategies. Open problems/sumarray/reference.py and paste your correct solution, reading standard input and writing standard output.
The array template already includes a working example reference: read n and an array, then print its sum. You can use it immediately to learn the workflow:
python -m gentest validate sumarray --seed 2026
python -m gentest generate sumarray --seed 2026 --zipInspect problems/sumarray/tests/Test001/sumarray.inp and sumarray.out. The first input is 1 followed by 0; its output is 0. The manifest is problems/sumarray/tests/manifest.json, and the archive is problems/sumarray/tests.zip.
To regenerate the set:
python -m gentest generate sumarray --seed 2026 --clean --zipOn systems where Python is named python3, use python3 in these commands. On Windows, py also works. GenTest executes Python references using the same interpreter you used to launch it.
For guided setup, run python -m gentest with no arguments. It asks for the problem name, template, language, test count and number of subtasks. It creates the files and prints the next command; edit your constraints before generation.
flowchart LR
accTitle: First Test Set Workflow
accDescr: Create a problem, edit its Python config, supply a reference, generate and inspect tests.
create[Create problem] --> config[Edit config.py]
config --> reference[Paste reference solution]
reference --> generate[Generate with a seed]
generate --> inspect[Inspect inputs and outputs]
Working directly from the repository needs no installation. Optionally install the CLI into your current Python environment:
python -m pip install -e .
gentest doctor
gentest list-templatesgentest and python -m gentest expose the same commands. Run them from the project root, or put --root /path/to/project before the command.
| Command | Purpose |
|---|---|
new <problem> --template array --lang python |
Create a config and exactly one reference file; never overwrite an existing problem |
generate <problem> |
Generate and validate inputs, execute references, publish a complete test set |
validate <problem> |
Check config, reference settings/path, generated inputs, validators and duplicates; no compilation, execution or output files |
list-templates |
List available starter configs |
doctor |
Show Python, project root and whether g++ is available |
| Generation option | Meaning |
|---|---|
--seed 2026 |
Global integer seed; omitted seeds are generated and printed |
--clean |
Replace the previous tests/ after the whole run succeeds |
--zip |
Write tests.zip containing TestXXX/ directories and the manifest |
--tests 30 |
Override the total count; explicit strategy/subtask counts must still fit |
--verbose |
Show each test's metadata and full framework tracebacks on errors |
validate also accepts --seed, --tests and --verbose. new accepts --tests to set the initial count. Problem names use letters, digits, underscores and hyphens and must be valid on Windows.
Existing output requires --clean. A failed run stops at the first failed testcase, publishes no new test set, and leaves a previous successful set intact. That previous set's manifest still describes its original seed. Successful regeneration without --zip removes an older tests.zip to avoid a stale archive. Compilation uses a temporary build directory, and each output file is written atomically only after successful reference execution and any cross-check.
GenTest/
gentest/
cli.py # Commands and beginner wizard
config.py # Python config loading and distribution checks
engine.py # Seeds, generation, validation, manifests and exports
runner.py # Checked Python/C++ execution
utils.py # Generation and formatting helpers
templates/ # Editable Python starter configs
problems/
sumarray/
config.py
reference.py # Or reference.cpp, selected explicitly in config
brute.py # Optional second reference
tests/
manifest.json
Test001/
sumarray.inp
sumarray.out
tests.zip # Optional export
solutions/ # Optional legacy <problem>.py or <problem>.cpp
config_sample.py # Original legacy example
config_sample/ # Original legacy templates, including corrected graphs
genTest.py # Legacy batch generation launcher
genUltils.py # Compatibility import of gentest.utils
app.py # Alternate launcher for the same CLI/wizard
tests/ # Automated tests for GenTest itself
new creates problems/ as needed. A fresh checkout does not need a solutions/ directory. Generated test sets, archives, bytecode and build artifacts are ignored by Git; configs and reference source files can be committed.
Use any functions, imports, loops or custom data structures needed by the problem. The engine does not interpret problem statements. Configs execute as Python code in your environment.
The modern generator returns the entire input as a string:
def generate(ctx):
n = ctx.rng.randint(1, 100)
return f"{n}\n"| Context field | Meaning |
|---|---|
ctx.test_id |
1-based test ID, used in Test001 etc. |
ctx.subtask |
1-based subtask index |
ctx.strategy |
One of the six strategy names below |
ctx.seed |
Deterministic per-test seed |
ctx.rng |
An independent random.Random for this testcase |
ctx.strategy_index |
1-based occurrence of this strategy across the run |
problemName determines .inp/.out filenames. totalOfTests determines how many test directories are created. solution selects the reference explicitly. Relative solution paths are resolved from the directory containing config.py, including paths with spaces.
subtasks = [30, 70] allocates 30% of the test count to subtask 1 and 70% to subtask 2. These are test allocation percentages; GenTest does not assign judge scores. Percentages must be finite, nonnegative and sum to 100.
Allocation uses cumulative floors: with 7 tests and [50, 50], subtask 1 gets 3 tests and subtask 2 gets 4. A small percentage can receive no tests. Zero-percent subtasks are skipped correctly. Use enough tests to cover each intended subtask.
Choose constraints inside the generator, for example maxN = [100, 100000] and limit = maxN[ctx.subtask - 1]. The wizard initially assigns nearly equal percentages and identical limits; customize both.
For exact counts, add subtask_counts = [3, 7] with totalOfTests = 10. Keep one count for each entry in subtasks; counts must sum to the total. This overrides percentage rounding while the percentage list must still be valid.
| Strategy | What you control |
|---|---|
fixed |
Hand-written regression or sample inputs in fixed_cases, or custom fixed logic |
boundary |
Minimum/maximum sizes and values, empty structures when permitted, endpoints |
special |
All equal, sorted, periodic, disconnected, star-shaped or other meaningful structures |
random |
Broad randomized coverage with problem-specific distributions |
adversarial |
Inputs aimed at likely wrong assumptions or poor complexity |
stress |
Maximum constraints for time and memory pressure |
Strategies are labels and dispatch controls, not automatic test-quality guarantees. Your generator decides what each means.
fixed_cases = ["1\n0\n", "2\n-1 1\n"]
strategies = {
"fixed": 2,
"boundary": 2,
"special": 2,
"adversarial": 2,
"random": "remaining",
"stress": 2,
}Groups run in dictionary order. Only one group may use "remaining"; it fills whatever is left after all numeric counts. Fixed inputs bypass generate, but still pass through validation and duplicate checking. If fixed_cases is nonempty, its length must equal the fixed count. With no explicit strategy settings, fixed inputs come first and random fills the rest.
The subtask and strategy sequences are allocated independently. To give each subtask its own deliberate strategy mix, use test_plan instead of strategies:
totalOfTests = 10
subtasks = [50, 50]
test_plan = [
{"subtask": 1, "strategy": "boundary", "count": 1},
{"subtask": 1, "strategy": "random", "count": 4},
{"subtask": 2, "strategy": "adversarial", "count": 2},
{"subtask": 2, "strategy": "stress", "count": 3},
]Plan counts must match both the total and each subtask's allocation. Remove unused fixed_cases when replacing a template's strategy plan. If an override such as --tests 3 cannot fit the configured groups, GenTest reports the inconsistency; edit the config or create a smaller starter with new --tests 3.
Every run prints its global seed before loading the config. Each test seed is the first 8 bytes, interpreted as an unsigned big-endian integer, of SHA-256 over the ASCII string global_seed:test_id.
The manifest records the problem, generator version, global seed, number of tests, and each test's ID, subtask, strategy, per-test seed and SHA-256 input hash. Inputs are UTF-8 bytes, written without platform newline conversion.
Keep the config, GenTest/Python versions and seed together for reproducibility. Use ctx.rng for all randomness; helpers accept rng=ctx.rng. Avoid time, OS randomness, unordered iteration, and mutable state shared between testcases. Imported third-party RNGs and external data sources remain under your control. The engine seeds standard-library global random separately for each testcase to support legacy configs, then restores its previous state. It also seeds global random while loading the config.
Regenerate a whole set with the manifest's global seed, not a reported per-test seed:
python -m gentest generate sumarray --seed 2026 --cleanFor an individual testcase, use its manifest entry and rebuild the context. Load the config with the original global seed so any import-time standard-library randomness also matches:
import json
import random
from pathlib import Path
from gentest.config import load_config
from gentest.engine import Context, input_content
folder = Path("problems/sumarray")
manifest = json.loads((folder / "tests/manifest.json").read_text())
item = manifest["tests"][6] # Test007
config = load_config(folder / "config.py", seed=manifest["global_seed"])
index = sum(t["strategy"] == item["strategy"] for t in manifest["tests"][:7])
ctx = Context(item["test_id"], item["subtask"], item["strategy"],
item["seed"], random.Random(item["seed"]), index)
print(input_content(config, ctx), end="")The same construction works for a failed testcase using its printed metadata and the strategy occurrence from your plan. Generation is sequential; configs should not start threads that depend on global random.
Each starter contains comments on constraints, boundaries, random coverage, special/adversarial structures and stress cases. Edit them for your actual problem. All templates except array create an intentionally failing reference placeholder, so an unfinished solution cannot silently produce empty answer files.
| CLI template | Input and design examples |
|---|---|
single-number |
One integer; minimum, maximum, powers of two, near-maximum values |
array |
n and integers; singleton, equal values, alternating extremes, maximum length |
string |
Lowercase string; singleton, repeated/periodic patterns, long common prefix |
multiple-test-cases |
T arrays; tiny cases, repeated cases, mixed sizes, maximum total size |
matrix |
Integer matrix; single cell, narrow matrix, zero and alternating matrices |
grid |
./# grid; open, blocked, checkerboard, maximum area |
permutation |
Permutation of 1..n; identity, reverse, displaced minimum, shuffled |
intervals |
Inclusive ranges; points, full spans, nested and random overlaps |
queries |
Point updates/range queries; endpoint indices, repeated updates, large workloads |
tree |
Undirected tree; singleton, path, star, random parents |
undirected-graph |
Simple graph; no edges, paths, random density, complete graph |
directed-graph |
Simple directed graph; paths, a cycle, random edges, all ordered pairs |
weighted-graph |
Simple weighted graph; endpoint uniqueness, low/high weights, dense stress |
dag |
Acyclic directed graph; chains, shuffled topological order, complete DAG |
The task is to sum an array. Save this as problems/sumarray/config.py after creating the problem:
from gentest import utils
problemName = "sumarray"
totalOfTests = 12
subtasks = [50, 50]
maxN = [20, 10000]
solution = {"path": "reference.py", "language": "python", "timeout": 2.0}
fixed_cases = ["1\n0\n", "2\n-1000000 1000000\n"]
strategies = {"fixed": 2, "boundary": 2, "special": 1,
"adversarial": 1, "random": "remaining", "stress": 2}
duplicate_policy = "warn"
def generate(ctx):
limit = maxN[ctx.subtask - 1]
n = limit if ctx.strategy in ("boundary", "stress") else ctx.rng.randint(1, limit)
if ctx.strategy == "boundary":
a = [10**6 if ctx.strategy_index % 2 else -10**6] * n
elif ctx.strategy == "special":
a = [0] * n
elif ctx.strategy == "adversarial":
a = [(-1)**i * 10**6 for i in range(n)]
else:
a = utils.integers(n, -10**6, 10**6, rng=ctx.rng)
return utils.toStringLines([n, a]) + "\n"
def validateInput(content, ctx):
n, *a = map(int, content.split())
if len(a) != n or not 1 <= n <= maxN[ctx.subtask - 1]:
raise ValueError("Wrong array length for this subtask")
if any(abs(x) > 10**6 for x in a):
raise ValueError("Array element out of range")The complete reference.py:
import sys
n, *a = map(int, sys.stdin.buffer.read().split())
print(sum(a))Run python -m gentest generate sumarray --seed 2026 --clean. Large values can make the sum exceed a 32-bit integer; a C++ reference should use long long.
This example asks for the number of connected components in a simple undirected graph. Create the files with:
python -m gentest new components --template undirected-graph --lang pythonReplace problems/components/config.py with:
from gentest import utils
problemName = "components"
totalOfTests = 10
subtasks = [100]
solution = {"path": "reference.py", "language": "python", "timeout": 2.0}
fixed_cases = ["1 0\n", "4 2\n1 2\n3 4\n"]
strategies = {"fixed": 2, "boundary": 1, "special": 1,
"adversarial": 1, "random": "remaining", "stress": 1}
def generate(ctx):
n = 100 if ctx.strategy == "stress" else ctx.rng.randint(2, 30)
if ctx.strategy == "boundary":
edges = []
elif ctx.strategy == "special":
edges = utils.tree(n, shape="star", rng=ctx.rng)
elif ctx.strategy == "adversarial":
edges = [(u, u + 1) for u in range(1, n - 1)] # One isolated vertex.
else:
m = n * (n - 1) // 2 if ctx.strategy == "stress" else ctx.rng.randint(0, n * (n - 1) // 2)
edges = utils.graph(n, m, rng=ctx.rng)
return utils.toStringLines([f"{n} {len(edges)}", *edges]) + "\n"
def validateInput(content, ctx):
rows = [list(map(int, row.split())) for row in content.splitlines()]
n, m = rows[0]
edges = rows[1:]
if not 1 <= n <= 100 or len(edges) != m:
raise ValueError("Invalid graph size")
pairs = set()
for u, v in edges:
if not 1 <= u <= n or not 1 <= v <= n or u == v:
raise ValueError("Invalid endpoint or self-loop")
pairs.add(tuple(sorted((u, v))))
if len(pairs) != m:
raise ValueError("Parallel edges are not allowed")The complete reference.py uses a disjoint-set structure:
import sys
it = iter(map(int, sys.stdin.buffer.read().split()))
n, m = next(it), next(it)
parent = list(range(n + 1))
def find(x):
while parent[x] != x:
parent[x] = parent[parent[x]]
x = parent[x]
return x
for _ in range(m):
u, v = next(it), next(it)
parent[find(u)] = find(v)
print(len({find(v) for v in range(1, n + 1)}))Run python -m gentest generate components --seed 42. The fixed cases' outputs are 1 and 2.
Import from gentest import utils and pass rng=ctx.rng to random helpers. Bounds are inclusive, sizes are nonnegative, and graph vertices are numbered 1..n.
| Helpers | Purpose |
|---|---|
integer(low, high), integers(n, low, high) |
Integer or array; integers supports distinct=True, sort="asc"/"desc" |
permutation(n, start=1) |
Shuffled consecutive integers |
distinct_array, sorted_array(..., reverse=True) |
Unique or sorted/reversed data |
duplicate_array(n, low, high, unique=3) |
Many duplicates from a small sampled value pool |
string(n, alphabet="abc"), binary_string(n) |
Custom-alphabet or binary strings |
matrix(rows, cols, low, high), grid(rows, cols, alphabet=".#") |
Integer matrices or character rows |
intervals(n, low, high) |
Inclusive (left, right) pairs with left <= right |
queries(n, factories) |
Select a callable per query; each factory takes an RNG and returns your row |
tree(n, shape="random") |
Random-parent, path or star tree |
forest(n, components=3) |
Acyclic forest with exactly the requested component count |
graph(n, m) |
Configurable graph helper |
undirected_graph, directed_graph, connected_graph, weighted_graph |
Convenience graph variants |
dag(n, m, connected=False) |
DAG oriented along a shuffled topological order |
toStringLines, toStringSpace, toStringLinesAll |
Row formatting, space formatting, or recursive line flattening |
graph options are connected, directed, weighted, allow_self_loops, allow_parallel_edges, weight_range=(1, 100), and rng. All boolean options default to false. Directed connected=True means weakly connected, not strongly connected. Trees/forests/DAGs also accept weighted and weight_range.
Simple graphs sample edge ranks without replacement, using O(n + m) memory rather than constructing all O(n²) potential pairs. Weighted edge uniqueness depends on endpoints, not weights. Connected generation first creates a spanning tree. These helpers do not promise a uniform distribution over all connected graphs or labelled trees. Impossible edge counts fail immediately.
An optional validateInput(content, ctx) runs before either reference. Return None/True for success, return False to reject, or raise ValueError with a specific explanation. It also runs for fixed inputs and during validate.
Set duplicate_policy = "warn" (default), "error", or "allow". Duplicate detection compares SHA-256 input hashes across the whole run. Warnings identify the current and earlier test IDs; failures include strategy, subtask and seed.
For C++ configuration:
solution = {
"path": "reference.cpp",
"language": "cpp",
"timeout": 2.0,
"cpp_standard": "c++17",
"compile_flags": ["-O2", "-Wall"],
"compiler": "g++", # Optional compiler command/path.
}C++ compiles once per run, with a 60-second compilation timeout. Each reference execution has its own timeout. Top-level timeout, cpp_standard and compile_flags provide defaults; solution dictionary entries override them. References execute with their source directory as the working directory and communicate through stdin/stdout. Outputs are captured as bytes; a crash or timeout never becomes an answer file.
To compare an optimized reference with a brute-force solution, add:
brute_solution = {"path": "brute.py", "language": "python", "timeout": 5.0}
cross_check_strategies = ["fixed", "boundary", "random"]
def cross_check(ctx):
return ctx.subtask == 1 # Keep the brute-force workload small.The default cross-check strategies are fixed, boundary and random. cross_check(ctx) is an optional additional filter; configure small constraints for selected tests. Comparison uses whitespace-separated output tokens, so spaces and line breaks may differ, but numeric formatting must agree. It is not a floating-point checker. The first mismatch reports the test ID, strategy, subtask, seed, input, and both outputs, and stops generation. Arbitrary problem-specific checks can also live in your own Python code.
Start with small hand-written cases whose answers you understand. Cover every legal boundary: one element, zero edges, extreme values, endpoint indices and the empty case if your statement allows it.
Add special structures that exercise different paths through a solution: equal values, duplicates, increasing/decreasing arrays, periodic strings, disconnected graphs, stars and chains. Add adversarial cases aimed at common mistakes: integer overflow, off-by-one indices, state retained between cases, assumptions about connectivity, deep recursion, or an algorithm that becomes quadratic on a particular ordering.
Use random cases to explore combinations you would not write manually. Vary sizes, value ranges and graph densities rather than choosing every parameter uniformly over its largest range. Check small random cases against a brute-force implementation.
Finish with maximum-constraint stress cases for time and memory. A useful suite combines all of these categories; a large pile of random inputs alone does not establish that a solution is correct. Review the manifest to confirm the intended strategy and subtask distribution actually ran.
Existing configs continue to support:
import random
import genUltils
problemName = "oldproblem"
totalOfTests = 20
subtasks = [50, 50]
def genInputContent(testID, curSubtask):
return str(random.randint(1, 100))testID and curSubtask remain 1-based. The old helper names and call signatures remain in genUltils; new code can use gentest.utils. Date generation now handles December correctly, and distinct rounded floats/dates use bounded sampling instead of potentially endless retry loops. Graph samples now print one edge per line and reject repeated endpoint pairs regardless of weight.
For legacy configs without solution, GenTest looks in solutions/<problemName>.py and .cpp. Exactly one must exist. If both exist, explicitly select solution, or set solution_language = "cpp"/"python". No silent Python preference remains. If both generator functions exist, generate(ctx) takes precedence over genInputContent.
python genTest.py generates all repository problems/*/config.py in sorted order, replacing successful prior sets and exporting ZIPs as the old batch command did. It stops at the first failure. With arguments it forwards to the canonical CLI, for example python genTest.py generate oldproblem --seed 42 --clean.
python app.py opens the same beginner wizard; it no longer bulk-creates both languages from name.txt. The original empty name.txt is retained as a legacy file. Old internal GUI/app classes and unchecked generation functions are not retained as public APIs. There is no GUI dependency. Output moves to problems/<folder>/tests/TestXXX/, eliminating the duplicated problem-name directory. The .inp and .out filenames retain the original style.
| Problem | What to check |
|---|---|
| Missing g++ | Run python -m gentest doctor. Install a C++ compiler providing g++ and add its executable directory to PATH, or select Python. |
| C++ compile error | Read the captured compiler stderr. Check syntax, included headers, cpp_standard and compile_flags. |
| Reference solution TLE | The error shows the test and timeout. Check for infinite loops and excessive complexity, or set an appropriate solution['timeout']. |
| Runtime error | Read the exit status and stderr. Reproduce using the reported test/global seed; check input parsing, invalid indexing and recursion depth. |
| Invalid config | Check Python syntax, positive total count, percentages summing to 100, and strategy/subtask counts matching the total. |
| Validator rejection | Read the validator's explanation with the test metadata. Correct the generator or the validator's interpretation of the statement. |
| Both Python and C++ files exist | Select the intended path/language in solution instead of relying on discovery. |
| Test directory already exists | Add --clean to replace the generated set after a successful run. |
| Module not found | Run from the repository root, or install with python -m pip install -e . in the interpreter you use. |
| Need a traceback | Repeat generate or validate with --verbose. |
Run the dependency-free test suite:
python -m unittest discover -s tests -vTests cover seeds/hashes, independent reproduction, subtask and strategy validation, duplicate policies, validators, transactional replacement, Python/C++ execution, failures/timeouts, cross-checking, graph invariants, CLI/wizard/templates and legacy configs. C++ execution tests skip if g++ is unavailable. CI runs Python 3.10–3.14 on Ubuntu and Windows.
Original authors: Khanh Tran and Ngat Do, Code Dream Programming Learning Center. Licensed under the MIT license.