I am a computational biologist and scientific workflow developer working at the intersection of AI for biology, single-cell genomics, computational neuroscience, and reproducible research software. I am currently a postdoctoral researcher at Hokkaido University, where I build, validate, document, and troubleshoot Python-based workflows for transcriptomics, spatial analysis, machine learning, and biological foundation models.
I am especially interested in making complex scientific analyses easier to test, reproduce, review, and extend. My work combines biological interpretation with explicit workflow interfaces, validation checkpoints, portable environments, contributor-facing documentation, and responsible AI-assisted development.
- Biological foundation models: Applying Geneformer to donor-aware single-cell classification, held-out evaluation, in silico perturbation, and biological interpretation.
- Scientific workflow curation: Organizing modular analyses with explicit inputs and outputs, configuration, validation checkpoints, limitations, and handoff documentation.
- AI-assisted software development: Using Codex and related tools for drafting, debugging, refactoring, and documentation with human review and scientific verification.
- Computational neuroscience: Contributing to multivariate neuroimaging research on language impairment after acute stroke using quantitative CT, voxel-based analysis, lesion mapping, and principal component analysis.
- Genomics and multi-omics: Developing reproducible analyses for single-cell, spatial, microbial, clinical, and environmental datasets.
A curated single-cell foundation-model workflow spanning donor-disjoint data preparation, Geneformer tokenization and fine-tuning, held-out evaluation, in silico perturbation, and spatial validation.
The repository documents workflow contracts, validation evidence, limitations, portable environments, migration procedures, troubleshooting, contributor guidance, and responsible AI-assisted development.
A reproducible multi-stage workflow using RepeatModeler2/RepeatMasker, STAR, TEcount, and DESeq2. The repository demonstrates explicit setup and input contracts, long-running process monitoring, recovery procedures, workflow consolidation, and separation of historical from canonical analyses.
- Define inputs, outputs, assumptions, and success criteria.
- Separate workflows into inspectable, restartable stages.
- Add schema, row-count, leakage, biological-control, and runtime checks where appropriate.
- Lock dependencies and record source revisions for reproducible execution.
- Document setup, limitations, troubleshooting, recovery, and contributor expectations.
- Use AI tools to accelerate development while retaining human review and verification.
Python · R · Bash · Linux · Jupyter · Git/GitHub · uv · Scanpy · pandas · NumPy · scikit-learn · Geneformer · transformer models · single-cell RNA-seq · spatial transcriptomics · computational neuroimaging · workflow validation · technical documentation · AI-assisted coding
Open to collaborating on:
- AI and foundation models for biological and neuroscience data.
- Reproducible scientific workflows and modular research software.
- Single-cell and spatial transcriptomics.
- Workflow validation, interoperability, documentation, and AI-agent instructions.
- Email:
k.dauyey.bio.nu [at] gmail [dot] com




