This is a repo for testing Txt2Onto for classification of human vs non-human host, using existing host labels as ground truth.
BioSample records for bacteria carry free-text metadata (host_comments, isolation_comments,
comments) and, for about a third of samples, a host_name. The host_name gives the label (human /
not human). The free text is embedded with PubMedBERT and a logistic regression is trained to predict
the label from the text alone.
- Stage 1 (done): supervised baseline. Validation F1 0.889. See reports/stage1_report.md.
- Stage 2 (not started): self-training on the samples that have no
host_name.
PLAN.md has the full design, the leakage checks and the progress log.
Tested on Linux (WSL2) with Python 3.14 and an NVIDIA GPU (RTX 4070, 8 GB). The pipeline expects CUDA; running on CPU has not been tested.
conda create -n txt2onto python=3.14
conda activate txt2onto
# torch first, from the CUDA wheel index
pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
# install this package (run from this folder)
pip install -e .requirements.txt lists the direct dependencies. requirements.lock.txt records the exact versions
of the environment the results were produced in. It is a reference, not an install file: it includes
the CUDA build of torch and a local package (onto_emb) that pip cannot fetch from PyPI.
Check the install:
pytest| File | In git | Notes |
|---|---|---|
data/human_labeled_biosamples.csv |
yes (20 MB) | Labeled table the pipeline reads. is_human: 1 human, -1 not human, 0 no host name. |
data exploration/entrez_mapping_checkpoint.json |
yes (1.5 MB) | BioSample to BioProject map, used to group samples for the split. |
data/splits/split.parquet |
yes | Frozen train/val/test split. Do not regenerate. |
data/bacteria_ids.json |
yes (5 MB) | NCBI taxon IDs under Bacteria (593,566 IDs), used to keep only bacterial samples. |
data/Whole_BioSample.parquet |
no | Output of src/data_retrieval.py. See "Rebuilding the dataset". |
data/bacteria_BioSample.csv |
no (669 MB) | Raw bacterial BioSample table that src/data_prep.py reads. |
data/processed/, data/embeddings/, artifacts/ |
no | Generated by the pipeline. |
The files in git are enough to run the pipeline. The raw tables are only needed to rebuild the dataset from scratch.
From this folder, with the environment active:
python -m txt2onto.run --stage allThis runs four stages in order. Each can also be run alone with --stage <name>.
| Stage | What it does | Time |
|---|---|---|
prep |
Cleans the labeled table and builds the text for each sample | seconds |
embed |
Embeds the texts with PubMedBERT (downloads the model on first run; writes a 473 MB cache) | about 40 s |
split |
Groups near-duplicate texts and BioProjects, reuses the frozen split, writes reports/leakage_report.json |
seconds |
stage1 |
Grid search, model fit, validation metrics and figures in reports/ |
about 3 min |
python -m txt2onto.run --stage stage1_figures redraws the figures without refitting.
Settings live in configs/pipeline.yaml.
This is optional. The labeled table is already in git.
1. Retrieve BioSample metadata. Run from this folder (the script uses relative paths):
python src/data_retrieval.pyThe script uses DuckDB to query the public OmicIDX BioSample
parquet file over HTTPS, so no NCBI account or key is needed. For every BioSample it collects the
sample taxon and gathers the submitter's free-text attributes into six columns: host_taxid,
host_health, host_name, host_comments, isolation_comments and comments. It prints how often
each column is filled and writes data/Whole_BioSample.parquet. The query scans the whole BioSample
archive, so expect a large download and high memory use.
The list of bacterial taxon IDs in data/bacteria_ids.json was built by
src/bacteria_selector.ipynb, which walks the NCBI taxonomy tree (nodes.dmp from the NCBI taxdump)
and keeps every taxon under Bacteria (taxon ID 2). The notebook has hard-coded paths to edit before it
will run.
2. Label the samples.
python src/data_prep.pyThis reads data/bacteria_BioSample.csv and writes data/human_labeled_biosamples.csv.
Not yet connected: step 1 writes
data/Whole_BioSample.parquet(all samples), while step 2 readsdata/bacteria_BioSample.csv(bacteria only). The retrieval script needs to save its bacteria-filtered table under that name before the two steps run end to end.
data exploration/data_prep.ipynb is the original exploration and labeling notebook. It has
hard-coded paths under /home/clayg/txt2onto/ that need editing before it will run. Its Entrez cell
reads credentials from the environment:
export NCBI_EMAIL="you@example.org"
export NCBI_API_KEY="your-key" # optional; raises the NCBI rate limitMuch of this repo was written by Claude (Anthropic's AI model, through Claude Code) working with Clay.
Written by Claude
PLAN.md,README.mdpyproject.toml,requirements.txt,requirements.lock.txt,.gitignoreconfigs/pipeline.yamlsrc/data_prep.py(a script version of the notebook's labeling logic)src/txt2onto/:__init__.py,config.py,data.py,embed.py,leakage.py,split.py,train.py,evaluate.py,plots.py,run.pytests/:test_data.py,test_split_leakage.py,test_evaluate.pyreports/stage1_report.md
Generated by running Claude-written code
data/human_labeled_biosamples.csv(regenerated bysrc/data_prep.py)data/splits/split.parquetreports/leakage_report.json,reports/stage1_metrics.json,reports/stage1_training.png,reports/stage1_validation.png
Written by Clay
src/data_retrieval.py(OmicIDX BioSample query)src/bacteria_selector.ipynband its outputdata/bacteria_ids.jsondata exploration/data_prep.ipynb(Claude edited two cells: the unlabeled-row selection, and the Entrez credentials)data exploration/entrez_mapping_checkpoint.json(output of the notebook)data/output.txt(host names matched by the human-term search, exported from the notebook)