Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Repository files navigation
PKUEEG directory index (2026-09-17) ================================== Data acquisition institution: SHRC, PKU. sub-XX/ses-dayY/eeg/ Raw continuous BrainVision EEG and sidecars participants.tsv / participants.json Participant metadata stimuli/audio/ Fifty shared story recordings stimuli/transcripts/ Story text stimuli/stimuli.tsv / stimuli.json Shared stimulus index derivatives/preproc_40hz/ 1-40 Hz EEG, sampled at 250 Hz derivatives/preproc_8hz/ 1-8 Hz EEG, sampled at 128 Hz derivatives/preprocessing_reports/ Availability of ICA/template processing records derivatives/stimulus_features/ Envelope, Mel, wav2vec2 and BERT; static-vector status derivatives/stimulus_annotations/ MFA, Pinyin and legacy word boundaries code/ Processing, feature extraction and structural/BIDS checks Release scope: behavioral comprehension accuracy and technical-validation analyses/results are deferred to a future dataset update. Historical/internal validation reports are maintained outside this release. Anonymous participant demographics and experimental event metadata remain included. The scope is also recorded in code/config/release_scope.json. Stories are allocated 17/16/17 across days (01-17, 18-33, 34-50). Raw recordings remain continuous by day; stories are not separate raw runs. Feature/annotation index paths are relative to the dataset root. Stimulus index audio/text paths remain relative to stimuli/. No eye-tracking data or measured individual electrode coordinates are implied. The directory reorganization on 2026-09-16 moved existing files in place, without copying the EEG or feature payloads or changing scientific values. A later, metadata-only correction on the same date changed channels.tsv unit labels from V to µV to match the BrainVision headers. It did not rescale any signal values. No missing arrays were invented. Earlier reports/checksums are preserved outside the public release directory. Publication readiness is assessed separately; a directory scaffold is not evidence of completeness. GitHub should host code/documentation, not this full data payload. Original study description (paths updated) ----------------------------------------- PKUEEG PKUEEG contains anonymized EEG from 25 healthy right-handed native Mandarin speakers during continuous speech perception and eyes-open rest. All participants completed three recording sessions. Day 1 contains stories 1-17, Day 2 stories 18-33, and Day 3 stories 34-50. Each session includes one approximately 300-s eyes-open resting interval. Ethics ------ The protocol was approved by the Institutional Review Board of Peking University (IRB00001052-25045). All participants provided written informed consent before data collection. The EEG data were anonymized before release. Acquisition ----------- Days 1 and 2 were recorded at 1000 Hz with a NeuroScan SynAmps2 system using 62 scalp EEG channels and two EOG channels. Day 3 was recorded at 1000 Hz with a NeuSen Wireless EEG/ERP system manufactured by Neuracle, with 59 scalp EEG channels and five auxiliary channels. Recording system, session order, and story set are confounded and should not be interpreted as independent experimental factors. Standard montage labels are provided, but no participant-specific electrode coordinates were collected. Events ------ Each events.tsv contains one selected interval per story and one rest interval. The original hardware marker stream remains in the BrainVision marker file. When duplicate or accidental markers were present, the released interval used the last adjacent marker pair whose separation was at least 150 s, matching the segmenting rule used for the preprocessing derivatives. Stimuli and behavior -------------------- Fifty Mandarin story audio files and matching transcripts are included. The audio was synthesized with the Guagua Audiobook platform. Stories 1-25 use the Dapiaoliang female preset; stories 26-50 use the Huazai male preset. Behavioral comprehension accuracy is planned for a future update. Story-level stimulus derivatives include a 100-Hz speech envelope; uniform 50-Hz Mel, BERT and Wav2Vec 2.0 representations; 12-layer word-level BERT vectors; and MFA TextGrid, Pinyin-Hanzi and corrected word timing annotations. Detailed axes, dimensions and provenance are recorded under derivatives/stimulus_features and derivatives/stimulus_annotations. Derivatives ----------- The release includes NPZ derivatives at 1-40 Hz/250 Hz and 1-8 Hz/128 Hz. Signal values are stored in microvolts. The Day-3 raw BrainVision data and 40-Hz NPZ segments were corrected by a factor of 1e-6 because the source BDF physical unit field was empty even though its physical range was expressed in microvolts. The 1-8-Hz derivative was regenerated uniformly from the corrected 40-Hz derivative. Current contents and reproducibility limits ------------------------------------------ This copy contains legacy derivatives; it is not the later repair workflow's output. The included EEG preprocessing entry point does not implement bad-segment rejection. Extreme amplitudes remain in some NPZ segments, and per-session ICA models, excluded-component lists and convergence records are not included. Raw Day-3 labels and derivative labels differ in capitalization; channel matching must account for this. The existing preprocessing code still has known case-sensitive montage and bad-channel matching limitations. The release contains 100-Hz envelope, 50-Hz Mel/wav2vec2/time-aligned BERT, and 12-layer word-level BERT arrays. word2vec_100hz is a status directory without arrays; a complete 100-Hz Mel/wav2vec2/BERT feature set is not included. The legacy word-level BERT script does not exactly reconstruct every tested released array because its terminal overlapping-window rule differs. Complete extraction pipelines for the other released features and an end-to-end reproducible MFA workflow are not included. code/model_downloads/ contains a status README, not pinned model-download workflows. The dependency list is incomplete. Internal validation reports are kept outside the release and are not certification of scientific reproducibility. These limitations were documented, not repaired, by the metadata/documentation correction. Acquisition reference and ground remain unavailable, and hardware filter settings are not documented; no unknown acquisition parameters have been inferred. Publishing and download links ----------------------------- See code/docs/PUBLISHING.md and code/docs/GITHUB_README.md for the GitHub/OpenNeuro release plan. No public URL or DOI has been assigned yet. Anonymous participant age and sex match the previously confirmed participant table; no names or initials are included.