Skip to content

Repository files navigation

Universal Brain

Universal Brain is a local-first executive runtime for coordinating models, tools, durable missions, permissions, recovery, and verification without assigning permanent authority to any single LLM.

It is an engineering project focused on one central question:

How can an AI-driven system execute useful work while keeping human intent, authority, state, evidence, and rollback structurally visible?

Start here

If you want to understand... Start with
What the system is This README + System Architecture
How model routing works Intelligence Fabric
How engineering missions recover Engineering Agency V5.3
Permissions and authority Threat Model + Alignment Invariants
What has actually been verified V5.3 Verification Record
How to reproduce target-machine validation V5.3 Validation Runbook

Architecture at a glance

Human request
     │
     ▼
Alignment / requirements
     │
     ▼
Deterministic Executive Kernel
     │
     ├── model routing / council
     ├── durable mission state
     ├── memory / world state
     ├── engineering workers
     └── Tool Gateway
              │
              ▼
      authorized external action
              │
              ▼
      verification + evidence
              │
              ▼
        canonical event/state

The design keeps canonical truth and authority outside the model. Models are leased for reasoning; deterministic control-plane components retain state, permissions, auditability, and recovery.

Current maturity

The repository contains implemented and tested engineering checkpoints through V5.3 validation tooling, but several real-environment claims remain deliberately gated until target-machine evidence exists. The project distinguishes mock/software verification from Windows, WSL2, Ollama, language-server, endurance, and physical-environment evidence.

That distinction is intentional: implemented is not treated as synonymous with proven in the target environment.

Why follow this project

Development is centered on concrete systems problems rather than model demos alone:

  • durable autonomous-mission recovery;
  • local/cloud model routing without surrendering canonical state;
  • explicit authority checks before consequential actions;
  • multi-repository engineering workflows;
  • semantic code intelligence and verification;
  • WSL2/Hyper-V isolation;
  • rollback and remediation;
  • target-machine endurance and evidence sealing.

If you are interested in AI agents, local AI, autonomous engineering systems, LLM orchestration, reliability, or safety-governed execution, this repository is where those experiments and verification records are published.


Governing Objective

Make silent misalignment structurally difficult, observable, and recoverable.

The system does not claim guaranteed access to unexpressed human intent. It instead establishes enforceable, machine-testable guarantees for explicit instructions:

  • Complete requirement and action traceability (REQ-ALN-001 to REQ-ALN-006);
  • Versioned constraints, assumptions, corrections, and commitments;
  • Fail-closed handling of high-impact ambiguity;
  • Centralized permission checks via a deny-by-default Tool Gateway before any consequential action;
  • Evidence-backed completion (primary measurements and tests over model assertions);
  • Model-independent state, memory, and continuity.

Approved Foundation Decisions

  1. Reversible Digital Actions: Version 1 may perform approved, reversible digital actions (A1) within declared scopes.
  2. Local-First Canonical Sovereignty: Canonical state, private data, and ledgers remain local-first; cloud models receive only scoped, minimized context.
  3. Neutral Human Safety: Human safety is not ranked by identity or relationship. Operator preference breaks ties only when credible human risks are materially equivalent.
  4. Universal Executive Brain (UEB): No LLM permanently sits at the top. The deterministic Executive Kernel owns canonical truth, while ephemeral Executive Models are leased for reasoning.
  5. Local Windows + WSL2/Hyper-V Execution (No Docker): Untrusted and model-generated execution boundaries use Windows-native isolation via dedicated WSL2 distributions or Hyper-V VMs without raw Docker dependencies. (Cloud geo-redundancy via Oracle Always Free is pending operator ratification under Decision D-007).

Milestone Implementation & Verification Status

Milestone Capability Area Implementation Status Verification Baseline
M1 Sovereign Kernel & Events Implemented (Exploratory) Unit & Master Gate passed (SQLite)
M2 Alignment Contracts & Ambiguity Implemented (Exploratory) Hardened in S1: Ambiguity defaults to MEDIUM
M3–M4 Executive Leases, Handoffs & EAP Mechanics implemented; Providers stubbed Mock verified; Real Ollama provider scheduled in S4
M5 Reversible Sandboxing & Workers Implemented (Exploratory) Hardened in S1: Request-bound capability & two-token worker auth
M6 Durable Persistence & Repositories Implemented (Exploratory) SQLite verified; PostgreSQL consolidation scheduled in S2
M7 Autonomous Missions & Coordination Implemented (Exploratory) Unit & Master Gate passed (SQLite)
M8 World Model & Situational Awareness Implemented (Exploratory) Unit & Master Gate passed; A2 actions locked in S1
S0 Working Copy Baseline & Manifest Completed & Verified 109/109 tests passed; Hash manifest generated
S0.5 Governance Reconciliation Active SAR-S0-S2 authorized; Draft status affirmed
S1 Authority, Worker Auth & A2 Lockdown Stabilized in current working slice Security regression suite passing in V3 checkpoint
IF-V4 Adaptive Cognitive Control / Multi-Transport Runtime Packaged & Verified Checkpoint 10 V4 tests; 45-test cognitive/integration gate; 18-test API/security gate passing
EA-V5 Autonomous Engineering Agency Runtime Packaged Foundation Checkpoint 14 dedicated V5 tests; historical 97-test engineering/intelligence/security gate
EA-V5.1 Production Engineering Integration Packaged & Verified Checkpoint 13 dedicated V5.1 tests; 102-test engineering/intelligence/security gate; 23-test API/core/security gate passing
EA-V5.2 Engineering Runtime Hardening Packaged & Cross-Platform Verified Checkpoint — target-machine evidence pending 19 dedicated V5.2 tests; 121-test engineering/intelligence/security gate; 32-test API/core/security gate passing
EA-V5.3 Real-Environment Validation & Endurance Validation harness implemented — target-machine execution pending Resumable target sessions, Ollama/pressure probes, authority-gated service workload, cross-process recovery, WSL2 execution probe, sealed evidence manifest

Master Documentation Index

1. Constitutional & Alignment Foundations

2. Architecture & Subsystem Specifications

  • System Architecture — Core modular monolith design, subsystem boundaries, and tech stack.
  • Executive Brain Specification — Executive Awareness Package (EAP), Action Commands, and Cognitive Handoff protocols.
  • Memory Architecture — Local-first storage and backup architecture.
  • Cost Control Policy — Hierarchical budgets, 3-tier exhaustion algorithm, and hard circuit breakers.
  • Operator Console Specification — React SPA layout and telemetry isolation.
  • Intelligence Fabric Specification — Model/route/transport separation, context compilation, routing, escalation, Council execution, interactive transports, telemetry, and safety invariants.
  • Engineering Agency V5.1 — Durable engineering recovery, requirement-scoped worktrees, impact-aware verification, causal cognition audit, local-resource admission, dependency diagnosis, and persistent service supervision.
  • Engineering Agency V5.2 — Tree-sitter/LSP semantic sensing, real stdio LSP runtime, opaque one-time ToolGateway isolation execution plans, durable worker fencing, multi-repository impact, transactional dependency remediation, verified merged-tree integration, authority-gated services, Operator Console telemetry, endurance harness, and target-evidence tooling.
  • Engineering Agency V5.3 — Resumable real-environment validation, three-model Ollama probes, optional concurrent pressure, disposable Git conflict proof, authority-gated service workload, cross-process recovery, real WSL2 execution probe, and offline-verifiable evidence bundles.
  • Platform Requirements — Traceable functional and non-functional requirements.
  • Threat Model & Security — 16 threat vectors, trust boundaries, and mitigations.

3. Architecture Decision Records (ADRs)

4. Testing & Verification


Current Status

The repository remains constitutionally governed by the existing stabilization authorization record, while the current working slice has advanced the provider layer into a real Intelligence Fabric / Multi-Transport Runtime. External live routes remain disabled until explicitly configured and authorized.

The V4 packaged checkpoint adds evidence-backed capability discovery, deterministic evaluation-driven routing, canonical/derived memory context adapters, Council disagreement adjudication, provider-session recovery, authorized Windows UIA support, and TaskDAG cognitive mission execution. The dedicated adaptive Fabric tests, Executive cognitive tests, M3/M4 master gate, cellular-monolith integration tests, and API/security regression gates pass in the recorded V4 verification run. Live external-provider/browser/desktop smoke testing remains explicitly operator- and environment-gated and is not claimed by this checkpoint. Persistence is now lazy at the FastAPI dependency boundary so importing/querying the control plane does not require a concrete database driver until durable persistence is actually used.

V5 adds an Autonomous Engineering Agency Runtime above the Intelligence Fabric: requirement-oriented hierarchical planning, cycle-checked DAG mutation, deterministic failure-classification/recovery insertion, a bounded cognition→ToolGateway→observation→repair loop, executable verification adapters, a local code-symbol/reference graph with test-impact retrieval, authority-gated Git worktree transactions, local Ollama resource profiles, independent worker completion gates, and a project Definition-of-Done auditor.

V5.1 advances that foundation into Production Engineering Integration: atomic self-verifying mission checkpoints with repository/worktree drift detection; incremental code-graph refresh and transitive impact analysis; impact-aware verification in the exact requirement worktree; requirement-scoped branch affinity and retry-safe verified integration; model-request/response digest events causally linked to ToolGateway evidence; router/runtime Ollama RAM/VRAM/concurrency admission; manifest/lockfile-aware dependency failure diagnosis; and durable long-running service state behind an authority-gated process backend. The V5.1 checkpoint records 13/13 dedicated tests, a 102/102 engineering/intelligence/security regression gate, a 23/23 API/core/security gate, and successful Python compilation. The unfiltered repository suite remains environment-blocked by unavailable pgvector; after excluding only that DB-model collection test, remaining non-passing cases are due unavailable aiosqlite in this sandbox.

V5.2 advances the runtime hardening layer with optional real Tree-sitter parsing, standard-LSP semantic/diagnostic/refactor adapters, a shell-free real stdio LSP process runtime, stale-diagnostic invalidation, explicit WSL2/Hyper-V isolation plans with fail-closed resource/network capability checks, and an opaque one-time isolation plan execution tool behind ToolGateway so model-authored host command strings are not accepted directly. It also includes durable generation-fenced worker leases, derived multi-repository dependency impact, dependency remediation with Git rollback, merged-tree verification before integration commit, a real ToolGateway-compatible persistent service process tool, read-only Engineering Agency API/Console telemetry, an atomic endurance harness, and a self-digested target-machine evidence collector. The recorded cross-platform gate is 19/19 dedicated V5.2 tests, 121/121 engineering/intelligence/security tests, 32/32 API/core/security tests, and successful Python compilation. The broader suite excluding the unavailable pgvector DB-model test reached 175 passed with 15 failures and 8 errors, all shown from unavailable aiosqlite in this sandbox.

V5.3 adds a dedicated real-environment validation harness rather than another architecture rewrite. It persists self-digested validation sessions with workspace-drift checks, probes the operator-selected Ollama models without retaining raw generations, optionally performs one bounded concurrent local-model pressure cycle, exercises disposable Git worktrees/conflict abort, runs a real bounded service through ToolGateway, proves checkpoint recovery across two Python processes, can execute one real WSL2 one-time isolation plan when the operator supplies the distro/path mapping, and seals target artifacts into an offline-verifiable SHA-256 manifest. The code/harness is cross-platform tested, but the Windows/Ollama/WSL2/multi-hour deployment evidence still must be produced on the operator machine.

V5.2 still does not claim demonstrated autonomous completion of a 50k–100k-line production system or production-proven Windows isolation. Real external language-server/Tree-sitter installation, WSL2/Hyper-V ToolGateway execution evidence, a >=2-hour real-repository endurance run, and dependency-complete full-suite verification remain target-machine proof gates.


License

All rights reserved. Dedicated to the operator until explicitly decided otherwise.

Intelligence Fabric / Multi-Transport Runtime

The provider-neutral universal_brain.intelligence layer separates model identity, access route, and runtime transport. V4 adds evidence-backed capability discovery, deterministic evaluation-driven routing, mission/event/semantic memory retrieval with provenance, independent Council disagreement adjudication, recoverable provider-session continuity, bounded browser attachments, a real optional Windows UIA chat driver, and adaptive TaskDAG cognitive execution that stops at Verification Engine / ToolGateway boundaries. New frontier models remain catalog/configuration data rather than Kernel architecture changes. See INTELLIGENCE_FABRIC_SPEC.md and ADR-0009.

About

Local-first multi-model executive runtime with authority gates, durable missions, model routing, recovery and evidence-backed verification.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages