Skip to content

About

Local, offline DFIR and security-analysis suite — wraps mature forensic tools as sensors, correlates their output across sources on one schema, and grounds findings in MITRE ATT&CK.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

EventHound

EventHound — cybersecurity analysis suite

Disclaimer — read this first.

This is a personal project and a proof of concept, not a finished product and not a commercial one — it never will be. It is published as is, with no warranty of any kind: it is still immature, it plausibly contains bugs, and there are certainly edge cases nobody has thought about yet. I have used it successfully on real cases, which is not the same as it being production software.

It was developed largely with AI, across several assistants (Claude Code, Mistral Vibe), under my guidance and review. That was half the point: I wanted to go deeper into what these tools can actually do and into the concepts around them, and I wanted that exploration to end up as something useful in my day-to-day work rather than as a demo.

The context is defensive analysis and authorized incident response only — not offensive tooling. Output may contain inaccuracies: validate against primary sources before acting on it operationally.

EventHound is a local, offline cybersecurity analysis suite that combines two main components:

  1. Analysis engine (analysis/) — ingestion of EVTX, PCAP, registry, MFT, THOR, CrowdStrike, osquery and generic logs, normalization into a common ECS schema, detection (Hayabusa/Sigma) plus the full Hayabusa toolbox (metrics, keyword/regex search, keyword pivots, base64 extraction), long-tail analytics (DuckDB), cross-source correlation, YARA matching, baseline/diffing, case management, threat-intel enrichment (Shodan/VT/ThreatFox, egress-gated) and compliance mapping (GDPR/NIS2/DORA). CLI-first; the web GUI (http://127.0.0.1:8700) is a thin interface over the engine. The knowledge that grounds an analysis — MITRE ATT&CK, GDPR/NIS2/DORA, ACN — lives as plain markdown under method/, read directly.

  2. Deterministic tools (tools/) — CVSS/EPSS scoring, compliance, Shodan/VT enrichment, workspace hygiene.

EventHound is built to run fully offline. Two surfaces do the whole job and either one is enough: 100% via the GUI and 100% via the terminal (CLI). There is no bundled LLM: reasoning about an analysis is done by an external agentic harness (Pi, Claude Code, …) through the analysis MCP server — analysis/analysis_mcp_server.py — which runs the pipeline and returns pseudonymized findings (§9). Analysis data never leaves the machine: the only outbound traffic is explicit, opt-in fetches of public resources — threat-intel / IoC lookups (Shodan, VirusTotal, ThreatFox) — each egress-gated and restricted to public indicators (§9).

External AI assistants are development tools, not part of the shipped product and not required to run it. If you want to use one with EventHound, you can — see Using an AI assistant with EventHound below. During development, real client data is handled under strict anonymization (§9); the product itself, being offline, has no cloud to send anything to.

Why this exists

Nothing here tries to replace anyone's work, and nothing here is claimed as original where it is not. The forensic tools this suite drives — Hayabusa, the Eric Zimmerman tools, Zeek, tshark, THOR — are other people's excellent work, wrapped and credited, never reimplemented.

The gap I kept running into is a different one: there are plenty of good security tools, and none of them talk to each other. Each one answers its own question in its own format, and the analyst is left doing the joining by hand — the same host spelled three ways, the same hash in two cases, the same hour reconstructed across four exports. The products that solve that properly exist, and they cost more than a small operation can justify.

So this is a wrapper: something that takes those heterogeneous outputs, puts them on one schema, and lets me ask what actually happened in a given case. That is the whole ambition — not a new detection engine, not a SIEM, not a product.

Why there is regulation in here

One other problem came from the same day job, and it explains the part of this repo that is not code.

Writing an incident report or a piece of technical documentation regularly means checking an obligation: does this qualify as notifiable, to whom, within how many hours, under which of GDPR / NIS2 / DORA. Re-reading a regulation from scratch every time is slow and, worse, error-prone. So the official texts and the ACN guidance are indexed locally and searchable (normative, acn), and tools/compliance/ maps the characteristics of an incident onto the obligations it triggers — as a starting point to verify, never as legal advice.

Vendor product documentation used to live here too — CrowdStrike Falcon and SonicWall, indexed for query syntax and product behaviour. It has been removed: both vendors ship an official MCP server, which answers those questions against the live product instead of against a snapshot of its manual, so that work belongs in a separate project. What stays in EventHound is the artifact side of those products: the CrowdStrike adapter still parses a detection export into the common schema, exactly like an EVTX or a PCAP.

The idea — tools as sensors, not the product

At the acquisition layer EventHound reinvents nothing: Hayabusa does the Sigma/ATT&CK matching, the Eric Zimmerman tools parse EVTX, MFT and registry, Zeek and tshark read the wire. These are existing, best-of-breed forensic tools — wrapped, not reimplemented (minimal code: reinvent nothing a tool already does well).

The value is the layer above and between them, which no single tool provides:

  • One schema. Each tool emits a different format; the adapters (analysis/adapters/) map them all onto a single ECS-subset schema. That common ground is what makes heterogeneous outputs comparable — without it you have seven silos.
  • Correlation. The DuckDB engine (analysis/analytics/) runs long-tail analytics (process stacking, rare parent-child, rare DNS, beaconing), builds temporal episodes and links indicators across sources — the same IP, user, host or hash seen in EVTX, PCAP and logs. This is analysis that operates over the tools' output; none of the wrapped binaries does it alone.
  • Interpretation. The local knowledge base (method/, ATT&CK map) grounds findings: the engine flags (T1558.003 on HOST-01), the knowledge base explains.

So EventHound is not a GUI over existing tools — it is a normalization and correlation layer that uses those tools as sensors. The tools are the probes; the work is putting them on one schema and surfacing the signal that only emerges by crossing them.

For the full component map, the end-to-end data flow and the offline boundary, see docs/architecture.md.

Scope of use

What it does

  • EventHound — EVTX/PCAP/log ingestion, ECS normalization, detection (Hayabusa/Sigma), long-tail analytics (DuckDB), cross-source correlation, YARA, baseline/diffing, case management, HTML reports, local GUI.
  • Knowledge base — MITRE ATT&CK (vendored as analysis/analytics/attack_map.json) and the GDPR/NIS2/DORA and ACN notes under method/, read directly as markdown.
  • Deterministic tools — CVSS/EPSS scoring, GDPR/NIS2/DORA compliance, Shodan/VT/ThreatFox enrichment.
  • Agentic analysis — the suite exposes itself to an external AI harness (Pi, Claude Code, …) through the analysis MCP server (analysis/analysis_mcp_server.py), which runs the pipeline and returns pseudonymized findings. External AI assistants are used to develop the suite too, under method/conventions.md; neither is needed to run it.

What it is NOT

  • Not a cloud SIEM: it operates locally and in batch on the supplied datasets, not in real time.
  • It does not run commands against remote systems: it proposes them, execution is up to the user (see method/conventions.md §12).
  • Not offensive: the context is exclusively defensive analysis and authorized incident response.
  • Not a cloud product: EventHound is self-contained and offline by design — not a hosted service, and it sends no analysis data anywhere (only opt-in public-indicator lookups reach the network).

Using an AI assistant with EventHound

Optional, and entirely your choice — the suite is complete without it. If you want one, clone the repo, start your assistant inside the repository, and it will pick up AGENTS.md: a briefing on what the tools are, how to call them, and the rules that matter when the data is real (client data never leaves the machine, security scores come from the deterministic oracle, commands are proposed and not executed). It works with any assistant that reads the AGENTS.md convention — Claude Code, Cursor, Codex, Gemini CLI.

That file also documents the optional .mcp.json that exposes the suite's own tools — analysis, CVSS/EPSS scoring, compliance mapping, threat-intel enrichment — as native MCP tools. It is not shipped enabled: what your agent launches should be your decision, not a default.

My own assistant configuration is deliberately not in this repository: it is personal workflow, not something you should have to download to use the suite.

The analysis MCP server

EventHound ships one MCP server — analysis/analysis_mcp_server.py — that exposes the analysis pipeline to an external agentic harness (Pi, Claude Code, Cursor, …). The reasoning lives in the agent; EventHound stays a deterministic local tool. It offers three tools: analyze (the whole pipeline over artifact paths), analyze_case (against a persistent case) and eid_lookup (the Windows Event ID vocabulary).

cd analysis && uv run python analysis_mcp_server.py    # stdio MCP server

Responses are pseudonymized (§9): hosts, users, internal IPs and domains are replaced with stable pseudonyms via data/pseudonym-map.md before the payload reaches the agent. If the map is empty the response carries a _privacy warning — treat that as a stop signal, not a formality. The same three tools plus the scoring, compliance and enrichment oracles can be enabled as native MCP tools; the .mcp.json block is in AGENTS.md, and nothing is enabled by default — what your agent launches is your decision.

Conventions

Comments across the codebase cite rules by number — (§9/§10) where client data is handled, (§6) where a security figure is produced, (§12) where a command is proposed rather than run. Those citations resolve in method/conventions.md: anonymization, source hierarchy, the git/privacy boundary, the posture on commands, hygiene. Reading it first makes the rest of the code read as intended. The engineering counterpart is method/minimal-code.md.

Development of EventHound is done with AI assistants under those same rules; their configuration is personal workflow and is deliberately not part of this repository.

Structure

path role nature
method/ how analysis is done and method knowledge: conventions.md (the § rules), minimal-code.md, security-instructions.md, anonymization.md, glossary.md, framework/ (MITRE/NIST/SANS/CIS), normative/ (GDPR/NIS2/DORA), fonti/ impersonal
docs/ analysis playbooks (analysis/), architecture.md, roadmap.md, brand/ impersonal
CHANGELOG.md what changed between versions, one screen; the reasoning stays in docs/roadmap.md impersonal
data/ real client data (anonymized) + pseudonym-map.md private
analysis/ EventHound: EVTX/PCAP/log ingestion, ECS normalization, detection (Hayabusa/Sigma), long-tail (DuckDB), correlation, YARA, baseline/diffing, case management, threat-intel enrichment, compliance, GUI (127.0.0.1:8700), HTML reports versioned (real data excluded)
tools/ deterministic tools: scoring (CVSS/EPSS), compliance (GDPR/NIS2/DORA), enrichment (Shodan/VT), workspace hygiene (check.sh, leak detection, injection scanning) versioned

Privacy and anonymization

Real data is sensitive. Work is done with stable pseudonyms (hosts, users, IPs, domains → HOST-01, USER-01, …); the real↔pseudonym mapping lives only in data/pseudonym-map.md (private). Full rules in method/anonymization.md.

Git/privacy model (the project uses git; data/ and the other sensitive directories are excluded from versioning via .gitignore):

  • Versioned: method/, docs/, analysis/ (code), tools/.
  • Excluded (.gitignore): data/ (client data), analysis/reports/, analysis/cases/, analysis/.tools/, Python environments.

Getting started (native — the recommended way)

The whole stack runs natively on macOS and Linux — no part of EventHound requires Docker.

git clone https://github.com/Stinocon/EventHound-Blue.git && cd EventHound-Blue
./setup.sh all              # install what's missing, start everything, then verify

Then open http://127.0.0.1:8700. That is enough to analyse evidence: the engine, the CLI and every source view work from here.

The installer downloads the forensic tools into the gitignored analysis/.tools/; nothing third-party is redistributed with this repository. If a download fails, if the machine has no outbound network, or if a version has to be pinned, docs/tools.md installs each component by hand and says what breaks without it. ./setup.sh doctor is the first thing to run either way: it reports what is present, what is missing, and the command that fixes each one.

Using it

Two interchangeable surfaces over the same engine — the GUI is a thin layer, never a shortcut around the CLI — and an external agentic harness on top of what they produce (see the paragraph at the top of this file for what it can and cannot do).

No evidence to hand? Run the demo first. It generates one coherent intrusion as eight source types — appliance logs, an EVTX detection timeline, a .reg, a THOR report, a CrowdStrike export, an osquery log, a YARA match and a capture — then ingests them through the ordinary adapters and produces the full analysis and report. Nothing in it is real: the estate is corp.example and documentation address space, so it is safe to show anyone. It is rebuilt from scratch on every run (--reset, the default), so the second run cannot fail on a batch the case already holds.

cd analysis
uv run python -m engine.run_demo                                     # generate, ingest, correlate, report
uv run python -m engine.run_demo --scenario triage-windows           # a Windows live triage (below)
uv run python -m engine.run_demo --artifacts-only                    # write the files, load them in the GUI
uv run python -m engine.run_demo --format markdown --level summary   # any run_report format/level

Two more datasets, both generated: a richer credential-theft and lateral-movement sample with each phase mapped to its ATT&CK technique, Event ID and Sigma rule (docs/samples.md), and a Windows live triage — one endpoint whose Security log was cleared before you arrived, so the only thing that still answers "what is happening now" is the state snapshot (docs/triage-windows.md).

The report lands in analysis/reports/. What the demo does not exercise is stated on every run — EVTX arrives as the JSONL Hayabusa emits rather than through the binary (pass --evtx-dir with real .evtx to include it), and a missing tshark or yara-python is reported, never quietly worked around. Loading the evidence one source at a time shows the point better than the all-at-once run: the correlation grows as it arrives, and the GUI's Cases view can step through it without the terminal.

GUI (http://127.0.0.1:8700): upload EVTX / PCAP / registry / logs / THOR reports per view, analyze, then read the Attack Map — entities and how they are linked, in kill-chain order, with a phase-by-phase account underneath and a click through to the timeline — plus the dashboard for cross-source correlation, and export from Report & Bundle. Load Demo Case in the Cases view fills it with the simulated incident described above, with nothing to upload. The in-app Help view is the per-view walkthrough; the correlation model (normalization → confidence → clusters) is documented in docs/analysis/correlation.md, and the hunting playbook in docs/analysis/threat-hunting-evtx.md.

What it looks like — the attack map (entities linked in kill-chain lanes, with the phase-by-phase account underneath) and the top of the report (kill chain, techniques, cross-source correlation), both over the sample intrusion:

Attack map

Report — kill chain, techniques and correlation

CLI (from analysis/, everything the GUI does):

uv run python -m engine.run_evtx sec.evtx --out report.md            # EVTX → ATT&CK detections
uv run python -m engine.run_logons sec.evtx                          # 4624/4625 logon view (lateral movement)
uv run python -m engine.run_pcap capture.pcap                        # flows, DNS, beaconing signals
uv run python -m engine.run_logs access.log --fmt access             # generic logs (syslog/access/jsonl/regex)
uv run python -m engine.run_mft '$MFT'                               # NTFS $MFT via MFTECmd
uv run python -m engine.run_analytics --evtx a.evtx --thor scan.txt  # long-tail + correlation, any source
uv run python -m engine.run_report --evtx a.evtx --out report.html   # visual report (also --format markdown|json)
uv run python -m engine.run_case new c1 --evtx sec.evtx               # persistent case (list/add/show/analyze/note/diff)
uv run python analysis_mcp_server.py                                  # MCP: analyze / analyze_case / eid_lookup

Cases — work that survives the session, and the default. Every analysis lands in a case (a directory under analysis/cases/ holding a DuckDB of the events plus notes) without being asked to: the correlation worth having happens between uploads. Sources accumulate across sessions, each upload reports what changed, and gateway/proxy/resolver addresses can be declared so they stop bridging every host to every other. A case is client data: analysis/cases/ is gitignored and refused by the leak guard. Full model and commands in analysis/README.md.

Bundles — reopening an analysis. One versioned JSON holding the conclusions and not the evidence, so a case reopens without re-running Hayabusa/tshark/EvtxECmd over data you may no longer have. Choose format Bundle (re-importable) in the GUI, or run_export / run_report --from-bundle on the command line. It carries the same real identifiers as any report: anonymize before sharing (§9). Detail in analysis/README.md.

Configuration and API keys

Third-party threat-intel services need a key. Set it once, from either surface — both read the same store, so a key entered in the GUI is already in place for the CLI and the MCP tools, and vice versa:

python tools/eventhound_config.py list            # masked status of every service + settings
python tools/eventhound_config.py set virustotal   # prompts without echoing the key
python tools/eventhound_config.py set-setting allow_egress true

In the GUI: Settings → API keys. The store is data/config.json — inside the project's private, gitignored area, written 0600, untouched by the uninstaller (so a reinstall finds its keys), and overridable with EVENTHOUND_CONFIG=<path>. An environment variable (VT_API_KEY, THREATFOX_API_KEY) always wins over the stored value, and the GUI says so when it does. Keys are never displayed again after saving — only their last four characters.

Storing a key does not enable network traffic: outbound lookups remain opt-in (allow_egress), and only public indicators are ever sent (§9/§15).

Docker — the alternative runtime

For non-macOS hosts or a reproducible deployment, the root docker-compose.yml brings up a single eventhound image with the engine, the GUI and every wrapped tool baked in.

Backup and moving to another machine

The code travels with the repo; four things do not (all gitignored) and are either recreated or backed up separately:

  • data/ — real client data + pseudonym-map.md. Private, never leaves the machine: back it up on your own terms, outside git.
  • analysis/.tools/ — the downloaded binaries (Hayabusa, the Eric Zimmerman tools, Zeek, tshark). Recreated by ./setup.sh install.
  • analysis/reports/ and analysis/cases/ — produced output and stored cases. Local-only by design; copy them yourself if you want them on another machine.
  • Python environments (.venv/) — recreated by uv sync.

Development and tests

One runner covers everything, and it is what "green" means here: the gate, the per-clone setup, the per-suite commands, and the fresh-clone check that catches what a tree with the installer already run cannot — all in docs/development.md.

Troubleshooting

Common problems and their answers are in docs/troubleshooting.md.

Status

docs/roadmap.md is the source of truth for what is built and what is open, and docs/architecture.md for how the pieces fit. This section used to be a third hand-maintained inventory of the same facts; three copies of one list drift, and then nobody knows which is current. In short:

  • EventHound — the engine is operational across ten sources, with correlation, the attack map, cases, reports and bundles, the local GUI, and a demo that runs the whole thing with no customer evidence at all.
  • Knowledge base — markdown under method/ that grounds analysis: framework/ (MITRE ATT&CK notes), normative/ (GDPR/NIS2/DORA), acn. ATT&CK technique→tactic is vendored in analysis/analytics/attack_map.json. Vendor product documentation is deliberately absent.

How mature is it, concretely

The disclaimer at the top of this file is not modesty, and here is the measurement behind it. The project's own release criterion is two consecutive clean adversarial review rounds. Four rounds have run; all four came back dirty — 11 defects, then 16, then a third round across four scopes, then a fourth round that found two of eleven sources had never worked on real data — and each round found defects created by the previous round's fixes. The counter is at zero. That is the honest state: the engine does real work on real evidence, and the rate at which review still finds things says it is not finished.

What that means in practice, if you are deciding whether to point this at something that matters:

  • Tested, lightly, on lab cases. The EVTX → Hayabusa/Sigma path, PCAP, generic logs, the correlation, the reports and the case model have all been run against lab inputs — the public EVTX-ATTACK-SAMPLES set, deterministic synthetic captures, and the simulated incident (engine.run_demo), which exercises eight sources end to end from files on disk. This is a proof of concept, not a hardened product: some function may be incomplete, inaccurate, or broken in a way no lab input has triggered.
  • Least-exercised: the osquery adapter was built against a documented schema and synthetic fixtures rather than a real client export, and say so in their own docstrings.
  • Not covered on a fresh clone: the EVTX/Hayabusa, MFT and registry-hive tests all SKIP without the binaries and a sample, so a green suite proves the adapters and the demo — not the headline source. -rs shows you which.
  • Known gaps carried deliberately, each with its reasoning in docs/roadmap.md: the kill chain derived only from ATT&CK evidence, which the endpoint sources are alone in carrying; and the host-artifact adapters (Prefetch, Amcache, LNK, SRUM, browser history) that wait on a real sample rather than on a guess.

Licence

EventHound is released under the MIT licence — see LICENSE. Use it, change it, ship it; keep the copyright notice, and expect no warranty (the disclaimer at the top of this file is not a formality — read it before pointing this at anything that matters).

What is not covered by that licence is listed in NOTICE.md, and the split is worth understanding before you redistribute anything:

  • Redistributed here: the Material Symbols icon paths inlined in the GUI (Google, Apache-2.0). That is the only piece of somebody else's work that arrives when you clone.
  • Not redistributed: every tool EventHound drives — Hayabusa, the Sigma rules it bundles, the Eric Zimmerman tools, Zeek, tshark. setup.sh install downloads them into the gitignored analysis/.tools/; each keeps its own licence and its authors' terms, which you accept by installing them. The Sigma community rules in particular are under the Detection Rule License, not MIT.
  • Not redistributed: the source material of the knowledge base (MITRE ATT&CK STIX, the official GDPR / NIS2 / DORA texts and ACN guidance). It is fetched or supplied locally and stays under its own terms; the regulatory answers this project produces are a starting point for verification, never legal advice.

Wrapping tools rather than reimplementing them is the design decision the whole project rests on (docs/roadmap.md); this is the licensing consequence of it, stated rather than left implicit.

About

Local, offline DFIR and security-analysis suite — wraps mature forensic tools as sensors, correlates their output across sources on one schema, and grounds findings in MITRE ATT&CK.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Contributors

Languages