Skip to content

About

Pipeline utilizing Txt2Onto 2 for host prediction (human/not human) from bacteria assembly metadata

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

humanpredict

This is a repo for testing Txt2Onto for classification of human vs non-human host, using existing host labels as ground truth.

BioSample records for bacteria carry free-text metadata (host_comments, isolation_comments, comments) and, for about a third of samples, a host_name. The host_name gives the label (human / not human). The free text is embedded with PubMedBERT and a logistic regression is trained to predict the label from the text alone.

  • Stage 1 (done): supervised baseline. Validation F1 0.889. See reports/stage1_report.md.
  • Stage 2 (not started): self-training on the samples that have no host_name.

PLAN.md has the full design, the leakage checks and the progress log.

Setup

Tested on Linux (WSL2) with Python 3.14 and an NVIDIA GPU (RTX 4070, 8 GB). The pipeline expects CUDA; running on CPU has not been tested.

conda create -n txt2onto python=3.14
conda activate txt2onto

# torch first, from the CUDA wheel index
pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

# install this package (run from this folder)
pip install -e .

requirements.txt lists the direct dependencies. requirements.lock.txt records the exact versions of the environment the results were produced in. It is a reference, not an install file: it includes the CUDA build of torch and a local package (onto_emb) that pip cannot fetch from PyPI.

Check the install:

pytest

Data

File In git Notes
data/human_labeled_biosamples.csv yes (20 MB) Labeled table the pipeline reads. is_human: 1 human, -1 not human, 0 no host name.
data exploration/entrez_mapping_checkpoint.json yes (1.5 MB) BioSample to BioProject map, used to group samples for the split.
data/splits/split.parquet yes Frozen train/val/test split. Do not regenerate.
data/bacteria_ids.json yes (5 MB) NCBI taxon IDs under Bacteria (593,566 IDs), used to keep only bacterial samples.
data/Whole_BioSample.parquet no Output of src/data_retrieval.py. See "Rebuilding the dataset".
data/bacteria_BioSample.csv no (669 MB) Raw bacterial BioSample table that src/data_prep.py reads.
data/processed/, data/embeddings/, artifacts/ no Generated by the pipeline.

The files in git are enough to run the pipeline. The raw tables are only needed to rebuild the dataset from scratch.

Run

From this folder, with the environment active:

python -m txt2onto.run --stage all

This runs four stages in order. Each can also be run alone with --stage <name>.

Stage What it does Time
prep Cleans the labeled table and builds the text for each sample seconds
embed Embeds the texts with PubMedBERT (downloads the model on first run; writes a 473 MB cache) about 40 s
split Groups near-duplicate texts and BioProjects, reuses the frozen split, writes reports/leakage_report.json seconds
stage1 Grid search, model fit, validation metrics and figures in reports/ about 3 min

python -m txt2onto.run --stage stage1_figures redraws the figures without refitting. Settings live in configs/pipeline.yaml.

Rebuilding the dataset

This is optional. The labeled table is already in git.

1. Retrieve BioSample metadata. Run from this folder (the script uses relative paths):

python src/data_retrieval.py

The script uses DuckDB to query the public OmicIDX BioSample parquet file over HTTPS, so no NCBI account or key is needed. For every BioSample it collects the sample taxon and gathers the submitter's free-text attributes into six columns: host_taxid, host_health, host_name, host_comments, isolation_comments and comments. It prints how often each column is filled and writes data/Whole_BioSample.parquet. The query scans the whole BioSample archive, so expect a large download and high memory use.

The list of bacterial taxon IDs in data/bacteria_ids.json was built by src/bacteria_selector.ipynb, which walks the NCBI taxonomy tree (nodes.dmp from the NCBI taxdump) and keeps every taxon under Bacteria (taxon ID 2). The notebook has hard-coded paths to edit before it will run.

2. Label the samples.

python src/data_prep.py

This reads data/bacteria_BioSample.csv and writes data/human_labeled_biosamples.csv.

Not yet connected: step 1 writes data/Whole_BioSample.parquet (all samples), while step 2 reads data/bacteria_BioSample.csv (bacteria only). The retrieval script needs to save its bacteria-filtered table under that name before the two steps run end to end.

Notebook

data exploration/data_prep.ipynb is the original exploration and labeling notebook. It has hard-coded paths under /home/clayg/txt2onto/ that need editing before it will run. Its Entrez cell reads credentials from the environment:

export NCBI_EMAIL="you@example.org"
export NCBI_API_KEY="your-key"   # optional; raises the NCBI rate limit

Who wrote what

Much of this repo was written by Claude (Anthropic's AI model, through Claude Code) working with Clay.

Written by Claude

  • PLAN.md, README.md
  • pyproject.toml, requirements.txt, requirements.lock.txt, .gitignore
  • configs/pipeline.yaml
  • src/data_prep.py (a script version of the notebook's labeling logic)
  • src/txt2onto/: __init__.py, config.py, data.py, embed.py, leakage.py, split.py, train.py, evaluate.py, plots.py, run.py
  • tests/: test_data.py, test_split_leakage.py, test_evaluate.py
  • reports/stage1_report.md

Generated by running Claude-written code

  • data/human_labeled_biosamples.csv (regenerated by src/data_prep.py)
  • data/splits/split.parquet
  • reports/leakage_report.json, reports/stage1_metrics.json, reports/stage1_training.png, reports/stage1_validation.png

Written by Clay

  • src/data_retrieval.py (OmicIDX BioSample query)
  • src/bacteria_selector.ipynb and its output data/bacteria_ids.json
  • data exploration/data_prep.ipynb (Claude edited two cells: the unlabeled-row selection, and the Entrez credentials)
  • data exploration/entrez_mapping_checkpoint.json (output of the notebook)
  • data/output.txt (host names matched by the human-term search, exported from the notebook)

About

Pipeline utilizing Txt2Onto 2 for host prediction (human/not human) from bacteria assembly metadata

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages