Skip to content

Command line

mhctools is one executable with subcommands. Run mhctools <command> --help for every option.

Command Does More
mhctools Predict for peptides or FASTA sequences below
mhctools ls List model artifacts and where they live getting models
mhctools fetch Download model weights and tool snapshots getting models
mhctools predictors Report which predictors can actually run (integrations is an alias) getting models
mhctools predict-table Append predictor scores to a CSV below
mhctools cleavage Per-bond peptidase evidence cleavage
mhctools vaccine-report Route-aware vaccine construct report vaccine reports
mhctools benchmark Assay-aware model evaluation benchmarks
mhctools mixtcrpred Score paired TCRs against a fixed target TCR predictors

The default command takes --mhc-predictor (one or more names, space or comma separated), --mhc-alleles or --mhc-alleles-file, and one input. The accepted names are listed by mhctools --help and in the CLI names column of the predictor matrix. TCR predictors have no name here because their input is a peptide plus a TCR.

Predict for peptides you supply

mhctools --sequence SIINFEKL SIINFEKLQ --mhc-predictor netmhc --mhc-alleles A0201

--sequence may be repeated and all occurrences accumulate. Or use --input-peptides-file for one peptide per line (blank lines ignored), or --input-fasta-file for protein sequences. Pick exactly one of the three.

Extract subsequences automatically

mhctools --sequence AAAQQQSIINFEKL --extract-subsequences \
    --mhc-peptide-lengths 8-10 --mhc-predictor mhcflurry --mhc-alleles A0201

Annotate a table (predict-table)

Reads a CSV, runs each requested predictor once, and appends one score column per predictor, choosing the best allele per row, while preserving every input column:

mhctools predict-table \
    --input benchmark.csv.bz2 \
    --peptide-column peptide \
    --alleles-column hla \
    --predictor netmhcpan42-ba:netmhcpan4.2.ba:affinity \
    --predictor netmhcpan42-el:netmhcpan4.2.el:score \
    --out benchmark.with_scores.csv.bz2

Each --predictor spec is NAME[:OUTPUT_COLUMN[:FIELD]], where FIELD is affinity, score, or percentile_rank. Lower is better for affinity and percentile_rank, higher for score.

A row may hold several alleles per cell (whitespace-, comma-, or semicolon-separated); the best one per peptide is chosen and recorded in a <OUTPUT_COLUMN>_best_allele provenance column. Missing or blank peptide/allele cells stay unscored. They are not coerced into a literal sequence or allele string and sent to a predictor.

Pass --predictor-info info.csv to also write a sidecar describing each column's score_field, units, and higher_is_better. Empty units means the field is dimensionless or predictor-specific.

Output conventions

CLI prediction tables follow one convention across every predictor. Plain peptide inputs get an empty source_sequence_name and offset 0, while FASTA and subsequence inputs keep their source and zero-based offset and are ordered by those coordinates. prediction_method_name is the exact CLI predictor name you selected, including version and mode. affinity is IC50 in nM, percentile_rank is a 0–100 percentile, and score stays predictor-specific. CSV floats are serialized with six significant digits.

Default stdout is streamed as tab-separated values, so an empty source name survives as an empty field and large tables don't need a second formatted copy in memory. A downstream closed pipe (| head, say) exits cleanly.