AiStockCN is one authenticated product for AI company research, evidence verification, quantitative signals, portfolios and controlled execution across more than 5,000 actively tracked US-listed equities and more than 5,000 China A-shares.
The integrated Research Copilot answers company questions, compares businesses and detects material filing changes while keeping original documents, financial facts, model inference and limitations clearly separated.
- Product: aistockcn.com
- US company research: aistockcn.com/us/research
- A-share company research: aistockcn.com/cn/research
- Product and operating guide: Research Copilot documentation
- Documentation index: docs/README.md
| Product capability | What it delivers |
|---|---|
| Market intelligence | Search and analyse 5,000+ US equities and 5,000+ China A-shares within one platform |
| Company research | Research US and A-share companies through market-scoped data, filings and financial evidence |
| Verifiable answers | Open the original document, page or SEC HTML passage behind a claim |
| Financial analysis | Use normalized SEC XBRL facts and deterministic calculations for revenue, profit, margins and cash flow |
| Company comparison | Compare two or three businesses through the same evidence and calculation workflow |
| Filing change detection | Identify added, removed, strengthened or weakened disclosures with both versions shown side by side |
The product entry connects Research, Verify, Quantify, Portfolio and Execute without presenting separate products for each market or workflow stage.
Desktop and mobile use the same product language, evidence contract and direct calls to action.
Select a company and review its market context and original filings. US research supports cited agent answers and standardized SEC facts. A-share ingestion prioritizes official SSE, SZSE/CNINFO and BSE disclosures and uses a separate Chinese retrieval profile.
Document claims retain their source identity throughout ingestion and retrieval. Native PDFs link to the original page. SEC HTML filings use honest passage locators and are never presented as paginated documents.
Run a consistent two- or three-company comparison using standardized financial facts, market calculations and retrieved filing evidence instead of independent free-form summaries.
Compare two annual reports to surface additions, deletions, strengthened language, weakened language and rewrites. Every proposed change includes both original passages, reproducible run parameters, history and a human-review decision.
Citation metadata is not generated by the language model. The server attaches document identity, filename, locator, page number and source URL from stored source records.
| Response layer | Meaning |
|---|---|
| Document evidence | Retrieved filing passages with verifiable source metadata |
| Financial and market evidence | Database facts and deterministic calculations with an as-of date and source locator |
| Model inference | Interpretation produced only from the supplied evidence and approved tool output |
| Limitations | Data scope, as-of dates and other context required to interpret the response correctly |
Verified evidence remains available independently from model-generated interpretation.
| Surface | Purpose |
|---|---|
/cn/overview, /us/overview |
Market overview and current product state |
/cn/research, /us/research |
Company evidence, filings, questions, comparison and filing changes |
/cn/quant, /us/quant |
Signals, methodology, walk-forward results and Explorer |
/cn/portfolio, /us/portfolio |
A-share holdings or US research and model baskets |
/cn/execution, /us/execution |
Controlled A-share execution or US readiness gates |
The global market switcher preserves the current stage. Each stage reports Live, In validation or Planned from the server-side capability gate. US execution does not connect a broker account or submit orders.
Administrative data is kept out of the customer research flow:
| Surface | Purpose |
|---|---|
/admin/research |
Filing coverage, ingestion state, failures and retrieval evaluation |
/us/models |
US model training and walk-forward governance |
/us/system-monitor |
US market-data pipeline health and run history |
/us/batch |
Scheduled US ingestion jobs |
All administrative routes enforce the administrator role on the server.
A research request is executed as a validated LangGraph state workflow:
- The authenticated frontend submits a company-scoped question.
- A typed LangGraph
plannode asks the configured Groq model for a schema-constrained JSON tool plan. - FastAPI removes unknown tools and enforces the server-side allow-list.
- The executor queries company data, standardized SEC facts, market history and document retrieval as required.
- PostgreSQL full-text and
pgvectorcandidates are fused with reciprocal-rank fusion. - A PyTorch cross-encoder reranks the retrieved passages.
- Deterministic tools calculate financial changes, returns and volatility.
- The configured model synthesizes only the supplied evidence and tool output, using server-assigned evidence IDs.
- A deterministic
validate_citationsnode rejects invented citation IDs and reports citation validity separately. - SSE lifecycle events keep the customer informed while node and tool timings are persisted for evaluation.
- The API returns evidence, inference, limitations, LangGraph trace and citation validation as structured data.
The production planner and synthesizer use Groq's OpenAI-compatible API with strict JSON schemas and a bounded timeout. If the provider is unavailable or rate-limited, the service falls back to a deterministic plan and verified evidence instead of inventing an answer.
flowchart LR
U["Authenticated user"] --> W["Next.js product"]
W --> C["Platform API"]
W --> M["US market API"]
W --> R["Research API"]
C --> CN["5,000+ A-shares<br/>market, models and portfolios"]
M --> US["5,000+ US equities<br/>prices, fundamentals and selections"]
R --> P["LangGraph state workflow<br/>plan, execute and validate"]
P --> X["SEC XBRL<br/>validated financial facts"]
P --> S["SEC and China official disclosures<br/>plus uploaded PDFs"]
P --> D["Deterministic financial<br/>and market calculations"]
P --> H["Hybrid retrieval<br/>FTS plus pgvector plus RRF"]
H --> RR["PyTorch cross-encoder<br/>market-specific reranker"]
P --> L["Groq GPT-OSS<br/>structured evidence synthesis"]
R --> Q["PostgreSQL work queue"]
Q --> K["Background ingestion workers"]
K --> V["Documents, chunks, vectors<br/>and source lineage"]
V --> H
R --> O["Structured logs, run history<br/>and retrieval evaluation"]
X --> P
US --> P
P --> A["Cited answer<br/>evidence plus inference plus limitations"]
A --> W
The research API and document workers are separate from the customer frontend, so ingestion and long-running analysis do not block ordinary navigation. US market services are isolated from the established A-share execution path while remaining part of the same customer product.
- Discover and synchronize 10-K, 10-Q and 8-K filings from the official SEC EDGAR archive.
- Discover and synchronize A-share reports through official SSE, SZSE/CNINFO and BSE disclosure identities.
- Upload annual reports and company filings as PDFs.
- Preserve SEC CIK, accession number, filing date, source URL and native locator type.
- Normalize SEC Company Facts into canonical annual and quarterly financial periods.
- Preserve exchange, announcement ID, report period, PDF checksum, source provider and native page numbers.
- Extract page-aware PDF text and honest SEC HTML passages.
- Chunk and embed documents in background workers backed by a PostgreSQL queue using
SKIP LOCKED. - Use separate English and Chinese embedding, FTS and PyTorch reranker profiles.
- Evaluate retrieval per market with Top-1 accuracy, mean reciprocal rank and a lexical baseline.
- Financial values and percentage changes are calculated by deterministic tools rather than recomputed by the LLM.
- Model plans and final output use validated structured schemas.
- Citation metadata is attached from database records after retrieval.
- Long-running responses emit SSE progress events and regular heartbeats.
- Bounded evidence context keeps inference efficient and predictable.
- Clear recovery actions support retryable requests.
- Verified evidence remains accessible independently from model synthesis.
- Ingestion runs, filing-change runs and human reviews retain durable history.
- Customer and administrative workflows use separate role-protected surfaces.
Research Copilot operates on the same live platform that supports:
- full-universe China A-share and US market-data workflows;
- feature engineering and inference snapshots;
- LightGBM training, scoring and model profiles;
- expanding-window walk-forward backtesting;
- atomic model activation through a PostgreSQL Model Registry;
- selection snapshots and operational monitoring;
- validation-gated paper-trading reconciliation;
- long-running market-data and reference-data orchestration.
| Area | Implementation |
|---|---|
| FastAPI and SSE streaming | apps/api/app/research_main.py, apps/api/app/routers/research.py |
| Multi-step agent and tool execution | apps/api/app/services/research.py |
| SEC filing discovery and sync | apps/api/app/services/research_sec.py |
| China official disclosure sync | apps/api/app/services/research_cn_disclosures.py |
| SEC XBRL normalization and calculations | apps/api/app/services/research_financials.py |
| Filing change detection and review | apps/api/app/services/research_filing_changes.py |
| PDF/HTML ingestion and workers | apps/api/app/services/research_documents.py, apps/api/app/research_worker.py |
| Hybrid RAG and pgvector | apps/api/app/services/research_retrieval.py |
| PyTorch reranking | apps/api/app/services/research_models.py |
| Retrieval evaluation | apps/api/app/services/research_evaluation.py |
| Unified customer UI | apps/web/app/cn/, apps/web/app/us/, apps/web/components/shell.tsx |
| US market API | apps/api/app/us_market_main.py, apps/api/app/services/us_market.py |
| Model Registry | apps/api/app/services/model_registry.py, scripts/create_model_registry.sql |
| US adjusted OHLCV and 5-day pipeline | scripts/backfill_us_daily_bars.py, scripts/train_us_5d_model.py |
| Container environment | docker-compose.yml, apps/api/Dockerfile, apps/web/Dockerfile |
| Tests | tests/test_research_service.py and the wider tests/ suite |
- Frontend: Next.js 15, React 19, TypeScript
- Backend: FastAPI, Uvicorn, Python 3.12
- AI and RAG: LangGraph, Groq GPT-OSS, structured tool planning, sentence-transformers, PyTorch, PostgreSQL FTS, pgvector
- Financial and ML: Pandas, PyArrow, LightGBM, scikit-learn
- Operations: Docker Compose, structured logging, retries, rate limiting and background workers
apps/api/ FastAPI APIs, agent, retrieval and ingestion workers
apps/web/ Authenticated customer and research interfaces
docs/ Product, architecture, results and operating documentation
run/ Safe example configuration and model profiles
scripts/ SQL migrations and operational runners
tests/ Unit and integration-style service tests
*.py / *.sh Data, model, backtest and trading workflows
AiStockCN uses real platform services rather than seeded sample data. A local integration environment requires:
- Docker and Docker Compose;
- PostgreSQL with the AiStockCN schema and
pgvector; - a Groq API key supplied through ignored runtime configuration;
- authentication values in
run/panel.envandrun/panel_users.json.
cp run/panel.env.example run/panel.env
cp run/panel_users.example.json run/panel_users.json
docker compose build research-api panel-web
docker compose up -d research-api research-worker research-coverage-worker us-market-api panel-web
docker compose ps research-api research-worker research-coverage-worker us-market-api panel-webThe Compose environment connects to the platform's existing external database and AI-service networks. See the detailed operating guide before provisioning a new machine.
docker compose config --quiet
docker compose run --rm --no-deps -v "$PWD/tests:/tests:ro" research-api \
python -m unittest discover -s /tests -p 'test_research_service.py'
npm --prefix apps/web run buildDatasets, uploaded documents, logs, model caches, runtime state and real credentials are excluded from Git. Safe configuration examples are provided in run/*.example; example credentials must never be used unchanged.
Research Copilot is an evidence-navigation and analysis system, not investment advice. Users should verify cited filings and financial data before making financial decisions.

