Persistent research memory for long-horizon mathematical reasoning.
Ansatz coordinates a main agent, parallel proof workers, a separate verifier, and a memory curator. It stores accepted proof records with their dependencies, tracks unfinished ideas in an exploration graph, and retrieves earlier research experience to guide subsequent attempts.
This anonymous release contains the Python implementation, agent contracts and skills, configuration templates, historical experiment summaries, and 13 research manuscripts: 10 PDFs and 3 Markdown proof write-ups. The implementation uses the cgm package, bin/cgm command, and CGM_* environment variables.
Paper: Continual Graph Memory for Mathematical Research Agents — link to be added after review.
Quick start · Research results · Architecture · Evidence and limitations · Configuration
- Preserves proof dependencies. A content-addressed fact graph records statements, proofs, and predecessors. Dependency traversal provides local proof context and supports revocation of dependent records.
- Maintains an exploration frontier. A typed graph connects goals, plans, obstacles, and findings. A curator updates annotations and promising directions as evidence changes.
- Separates proposal from admission. Workers propose facts; a separate verifier session reviews them before admission. Role-specific MCP tools constrain how each agent can read or change shared state.
- Reuses experience with explicit scope. Cross-project recall can be disabled, limited to the same batch, or enabled across projects. Recalled mathematical statements require local justification before becoming premises.
- Supports persistent runs. Project files, assignments, and worker state remain on disk; the CLI and dashboard expose progress across multiple reasoning rounds.
flowchart LR
P[Problem] --> M[Main agent]
M --> W[Parallel workers]
M <--> S[Strategy consultation]
W --> V[Separate verifier]
V -->|Accepted records| F[Fact graph]
F -->|Relevant dependencies| W
W --> E[Exploration graph]
C[Curator] <--> E
E --> M
X[Scoped research memory] -->|Recall for local justification| W
Here, accepted means accepted by the configured model-based review process. It does not imply a formal proof-assistant certificate or independent expert validation.
The following table summarizes claims made in the supplied manuscripts. All result files are preserved byte for byte, with neutral filenames. Scope and proof status are stated separately so that a result on one case is not mistaken for a solution to the full problem.
| Manuscript | Reported result | Scope and status |
|---|---|---|
| Erdős #149 | At maximum degree four, characterizes every strong clique of twenty edges as a component isomorphic to the two-vertex blow-up of a five-cycle. | Structural progress and colouring reductions; the degree-four strong-colouring conjecture is not resolved. |
| Erdős #289 | For every sufficiently large k, represents 1 as a reciprocal sum over exactly k disjoint, non-adjacent integer intervals, each containing at least two integers. |
Claimed resolution of the separated-interval formulation; the threshold is ineffective. |
| Erdős #302 | Candidate proof of f(N) = cN + o(N) for sets avoiding distinct reciprocal triples, with two characterizations of the limiting density. |
Candidate proof of a density limit; the numerical value of c remains undetermined. |
| Erdős #348 | Under universal quantifiers for both deletion clauses, classifies the possible pairs as (0,n) for n >= 1 and (1,n) for n >= 2. |
Claimed classification for indexed sequences with repeated values allowed; an existential second clause gives a different answer. |
| Erdős #488 | Gives an explicit finite construction claimed to violate the proposed density-of-multiples inequality, with density ratio greater than 256/67. |
Counterexample manuscript with two proof sketches; reports model review and small-scale numerical calibration. |
| Erdős #561 | Claims the star-forest size-Ramsey formula for two stars against an arbitrary star forest, plus an all-even diagonal-maxima case. | Partial result; some input and structural lemmas are not expanded in the supplied note. |
| Erdős #576 | Studies barriers to improving cube Turán bounds, with unbalanced-host results and conditional higher-dimensional reductions. | No improvement to the cube's 3/2 lower or 8/5 upper exponent is claimed; some proof dependencies are unrendered. |
| Erdős #644 | Claims f(k,7) <= ceil((6k+3)/7)+1 for k >= 150, with asymptotic coefficient 6/7. |
Partial bound toward the conjectured 3/4; key covering case trees remain to be fully checked. |
| Erdős #766 | Gives bounds and selected strict comparisons for minimum Turán numbers at prescribed vertex and edge counts. | Partial results; general strict monotonicity remains unresolved, and some proof dependencies are unrendered. |
| Erdős #811 | For every q >= 4, constructs arbitrarily large balanced binom(q,2)-colourings without a rainbow K_q. |
Claimed resolution of the clique case; the general graph question, including C_6, is not resolved here. |
| Erdős #933 | Claims unimodality of the independent-set sequence of every forest of order n >= 2^2017 + 2, with further log-concavity bounds. |
Finite reduction, not the all-order conjecture; smaller orders remain unresolved and bibliography placeholders remain. |
| Erdős #1063 | Establishes existence of the least binomial-divisibility defect and gives a super-polynomial lower bound and a refined subexponential upper bound. | Two-sided growth bounds; a sharp matching asymptotic is not claimed. |
| Jamison caterpillar conjecture | Presents an all-order argument that every tree maximizing mean subtree order at fixed order is a caterpillar. | Explicitly a proof candidate, pending independent mathematical verification. |
See the result index for precise conditions and manuscript locations. No new mathematical referee review was performed for this release. The #811 and Jamison PDFs describe companion verification programs or data that are absent from the supplied files; their reported computations are not reproduced by the harness tests.
The current paper source reports 10/10 outcomes on First Proof Second Batch. The supplied archival paper PDF and supporting run snapshot record an earlier 9/10 local-verifier outcome, with task 03 open and a different recorded model configuration. These are distinct records; this release does not establish that the newer claim has been reproduced from the older archive.
Historical component-study records are included under experiments/. They contain documented answer-route exposure in several ablation arms, which limits causal comparisons. The evidence notes explain the version and protocol differences. Benchmark outcomes, research-manuscript claims, and software tests are separate forms of evidence.
Use Python 3.10 or later. From the extracted repository root, the following example creates a local project and assignment, then inspects its state. It needs no third-party Python packages and launches no model calls.
python3 -m cgm.orchestration --help
(
demo_dir="$(mktemp -d)"
export CGM_AGENTS_ROOT="$demo_dir/projects"
export CGM_CONSULT_TRANSPORT=off
bash bin/cgm new demo --roles high:1
cp examples/project/PROBLEM.md "$CGM_AGENTS_ROOT/demo/PROBLEM.md"
bash bin/cgm assign demo/high --task 'Read PROBLEM.md and plan a proof.'
bash bin/cgm status demo --json
bash bin/cgm list --json
)The worker should remain in the created state, with no live process and zero completed rounds. The files under examples/project/ are illustrative hand-authored data, not a recorded proof-search result.
The bundled deployment scripts target Linux. They require Python 3.10+, Bash, curl, tar, and Git. The resident main agent and curator also need tmux and an installed, authenticated Claude Code CLI. Worker and verifier sessions require a configured Codex backend; strategy consultation uses the transport selected in your configuration.
-
Provision the local runtime and copy the configuration templates:
bash scripts/bootstrap.sh cp config/cgm.env.example config/cgm.env cp config/codex.env.example config/codex.env
-
Fill in your backend credentials, model identifiers, and consultation settings. Set up the selected backend using the commands in Getting started. Credentials belong in the ignored
config/*.envfiles. -
Start verification and check the deployment:
bash scripts/services.sh up verify bash scripts/doctor.sh
Bootstrap downloads dependencies. The health check can contact the configured model backend. These commands belong to live deployment, not the offline example above.
-
Open the main-agent session from the repository root:
claude
Follow the initialization workflow, supply the problem statement, and choose worker roles and stop conditions. After the project exists, start its curator with
bash examples/ops/curator-tmux.sh <project_name>. See the operating guide for the complete workflow and operations for service management.
LaTeX and a headless browser are optional output dependencies for paper and progress-report rendering. They are unnecessary for the offline scaffold example.
| Setting | Purpose |
|---|---|
CGM_CODEX_MODEL, CGM_CODEX_EFFORT |
Default model and reasoning effort for worker/verifier launchers; service-specific overrides are available. |
CGM_CONSULT_TRANSPORT |
Selects gpt_pro, claude_api, claude_code, or off. |
CGM_MEMORY_SCOPE |
off by default; batch restricts recall to matching batch labels; all enables cross-project recall. |
CGM_AGENTS_ROOT |
Directory containing project state. |
VERIFY_PORT, DASHBOARD_PORT |
Local verifier and dashboard ports; defaults are 8091 and 8099. |
Configuration defaults are deployment examples, not a declaration of the settings used for every reported experiment. See configuration and security and trust.
| Path | Contents |
|---|---|
cgm/ |
Memory stores, MCP gateway, verification, orchestration, execution, strategy, dashboard, and document utilities. |
agents/ |
Role contracts and worker/verifier skills. |
.claude/skills/ |
Main-agent workflows. |
bin/ and scripts/ |
CLI wrappers, deployment, and service management. |
config/ |
Configuration templates with placeholders. |
examples/ |
Toy project and operational examples. |
experiments/ |
Historical experiment summaries and their methodological caveats. |
results/ |
Thirteen supplied research manuscripts (10 PDFs, 3 Markdown files), their manifest, and an annotated result index. |
docs/ |
Architecture, usage, configuration, and release notes. |
SHA256SUMS |
Checksums for release-file integrity. |
Install the development extras in your own virtual environment, then run the existing tests:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'
python -m pytest cgm/ -qThe test suite checks software behavior using local fixtures and mocked transports. It does not prove the mathematical claims or reproduce the paper's benchmark. Release-specific validation is recorded in release notes.
Personal contact details, repository-owner links, deployment paths, machine-local settings, and private operator history have been removed or replaced in this release. Scholarly references and existing legal attribution are retained. The source-paper archive is kept separately, unchanged; its publication link will be added above when available.
The code retains its Apache 2.0 license and existing copyright notice. Bundling the research manuscripts does not assign them a new license. Citation metadata for this work will be added with the paper link after anonymous review.