VarSage is a bioinformatics decision-support platform for post-calling interpretation of small germline sequence variants. It is intended for exome or genome VCF/table outputs that have already passed through an upstream alignment, variant calling, and primary quality-control workflow. VarSage does not call variants from reads; it integrates existing variant calls with phenotype, inheritance, population, transcript, clinical-assertion, and computational evidence to produce an auditable prioritization dossier.
The application is organized around three analytical contexts: phenotype-driven rare-disease prioritization, couple-level carrier screening, and family-aware trio analysis. Across all contexts, VarSage treats the variant as a structured allele defined by chromosome, 1-based position, reference allele, alternate allele, selected sample genotype, and annotation fields. This deterministic coordinate model is deliberately kept separate from AI-generated summaries.
Clinical boundary: VarSage provides research and expert-review support. It does not establish a molecular diagnosis, does not issue a clinical laboratory classification, and does not replace orthogonal confirmation, segregation analysis, phase assessment, disease-specific ACMG/AMP review, or qualified clinical sign-out.
Rare Disease
Single-proband analysis that ranks candidate alleles by phenotype fit, molecular consequence, rarity, inheritance compatibility, quality, and provisional evidence.
Carrier Screening
Independent partner review followed by same-gene comparison for recessive reproductive-risk signals and scope-specific carrier findings.
Trio Analysis
Proband-first prioritization followed by parental allele-state comparison for possible de novo, inherited, and compound-heterozygous models.
VarSage follows a layered interpretation architecture. First, submitted phenotypes and variants are normalized into controlled representations. Next, transcript and coordinate annotations are attached while preserving original source fields. The resulting candidate set is filtered for assay scope, quality, benign assertions, and population frequency. Finally, retained variants are ranked using phenotype similarity, inheritance compatibility, molecular consequence, rarity, clinical assertions, computational predictors, and provisional ACMG/AMP-oriented evidence.
The final report is not just a score table. It is an audit trail: which sample was analyzed, which file type was detected, which variants were removed, which databases matched, which HPO terms were used, which inheritance model was plausible, and which evidence remains incomplete. This is essential in clinical bioinformatics because the same numeric rank can arise from very different biological and evidentiary situations.
Rare disease
Use when an affected proband has WES/WGS variants and the primary question is phenotypic plausibility of candidate alleles.
Carrier screening
Use when two partners are assessed for heterozygous variants in the same recessive disease gene or in configured carrier scopes.
Trio analysis
Use when maternal and/or paternal variant files can inform de novo, inherited, or biallelic hypotheses in the proband.
The rare-disease workflow is a single-proband prioritization pipeline for suspected Mendelian disease. Its purpose is to reduce a large variant set to a biologically interpretable shortlist while preserving the evidentiary chain behind each retained candidate. The workflow explicitly distinguishes review priority from pathogenicity classification and separates variant-level evidence from case-level causality.
Case intake. VarSage records patient metadata, selected genome build, sample identifier, sex, age, expected inheritance, family history, consanguinity, present HPO terms, explicitly absent HPO terms, clinical narrative, optional BED scope, and per-case AI permissions.
Phenotype preparation. Manual HPO identifiers are used as high-confidence anchors. Free text is mapped to HPO by deterministic term and synonym matching, and optionally by AI extraction constrained to valid HPO terms supported by the submitted description.
Variant parsing and normalization. VCF, VCF.gz, TSV, CSV, and delimited text are converted into a common allele model. Chromosome labels, 1-based positions, reference and alternate alleles, genotype, zygosity, depth, genotype quality, gene, transcript, consequence, and source annotations are standardized.
Transcript and coordinate annotation. Existing VEP CSQ, ANN, INFO, FORMAT, and table annotations are preserved. Jannovar can add transcript consequences. Local databases are merged by exact allele identity to attach frequency, clinical assertion, rsID, predictor, splice, or laboratory-specific evidence fields.
Pre-ranking filtration. The candidate set is reduced using benign/likely benign assertions, population-frequency thresholds, depth, genotype quality, VCF FILTER status, and BED interval overlap. Rescue logic keeps selected rare high-impact or clinically asserted variants visible for expert review.
Phenotype-gene similarity. Candidate genes are compared with HPO-derived gene and disease profiles. Ontology ancestry is used to recognize specific-to-broad phenotype relationships while avoiding unsupported inference of co-symptoms.
Inheritance assessment. VarSage evaluates compatibility between genotype state, sex, expected inheritance, family history, known gene inheritance modes, and disease models such as autosomal dominant, autosomal recessive, X-linked, homozygous recessive, and possible compound heterozygosity.
Evidence scoring and ranking. Integrated ranking combines phenotype match, inheritance support, molecular severity, rarity, call quality, clinical assertions, predictors, and provisional ACMG/AMP evidence. The rank orders review priority, not diagnostic certainty.
Interpretation support. Reports include provisional evidence, candidate rationale, literature context where available, secondary-finding and important-heterozygous sections, missing evidence statements, and optional AI-assisted review.
Export and audit. Results are written as browser views and downloadable reports so each prioritization decision can be inspected, reproduced, and discussed in a multidisciplinary review setting.
Case decision spectrum
After ranking and optional synthesis, VarSage adds a cautious case-level decision spectrum. This summarizes whether the retained candidate set contains no positive candidate, an inconclusive candidate, a positive candidate requiring confirmation, or an externally confirmed cause. The decision is derived from structured evidence and report-review signals; it is not a molecular diagnosis.
The same evidence spectrum used in rare-disease result and report views. Green is the cautious no-positive-candidate end; red is reserved for expert and laboratory confirmation.
No positive candidate identified
Assigned when no retained candidate is available, or when the retained candidates do not contain pathogenic/likely-pathogenic or primary report-review evidence. This is not a statement that a genetic diagnosis is excluded.
Inconclusive candidate finding
Assigned when candidates remain important for review, but the current evidence is not strong enough for a cautious positive decision. The result should remain inconclusive or candidate-only until the uncertainty is resolved.
Positive candidate finding
Assigned when a prioritized candidate has a pathogenic/likely-pathogenic provisional classification, a matching ClinVar-like assertion, or a primary report-review signal. Expert review, orthogonal confirmation, and relevant segregation or phase checks are still required.
Confirmed cause
This is a reporting endpoint for human interpretation, not an automatic VarSage decision. The tool never assigns it; it requires qualified expert review and laboratory confirmation where appropriate.
How the decision is taken: VarSage starts from the ranked candidate set. When case synthesis supplies reportable or decision-making candidate IDs, the decision check is narrowed to those candidates; otherwise, all ranked candidates are considered. It then checks pathogenic/likely-pathogenic evidence and primary report-review hints before checking whether a high-priority uncertain candidate still needs review. If none apply, it returns the cautious no-positive-candidate state. Candidate IDs supporting the decision are preserved for audit.
Ranked candidates
The main table orders variants by a composite score that combines phenotype fit, inheritance support, consequence, rarity, and quality.
Evidence details
Each row can expand to show annotation, filtering, phenotype overlap, literature links, and provisional ACMG evidence.
Report outputs
The workflow exports HTML, TSV, and Microsoft word results so the shortlist can be reviewed and shared outside the app.
The carrier-screening workflow is built for reproductive-risk review in a couple. It analyzes each partner independently, applies carrier-scope filtering, and then performs a gene-level comparison between retained heterozygous findings. This design prevents partner-specific evidence from being conflated while still making shared-gene signals immediately visible.
Couple intake. VarSage records partner identifiers, sample IDs, genome build, relationship and ancestry notes, reproductive concern notes, carrier panel settings, optional secondary-finding scope, and optional BED interval restriction.
Independent partner parsing. Each partner file is normalized separately. Genotype, zygosity, depth, quality, gene, consequence, clinical assertion, and predictor evidence remain attached to the correct individual.
Carrier-scope matching. Retained variants are assigned to explanatory scopes: configured carrier-panel genes, phenotype-concern genes derived from HPO mapping of notes, ACMG secondary-finding genes when enabled, and X-chromosome review where relevant.
Heterozygous carrier filter. The core carrier model focuses on heterozygous variants. Common, low-quality, benign, or out-of-scope variants are removed before couple-level comparison.
Partner summaries. Each partner receives an auditable retained-variant list with gene, variant, category, zygosity, allele frequency, evidence score, and rationale.
Same-gene comparison. The final couple step flags genes where both partners carry retained heterozygous variants. This is a review alert for recessive disease risk, not an automated fetal-risk prediction.
Risk wording and review notes. The report separates evidence from interpretation and leaves disease mechanism, phase, classification, residual risk, ancestry context, and counseling decisions to the qualified reviewer.
Outputs. VarSage produces partner-specific tables, same-gene warning tables, downloadable TSV files, HTML reports, and Word reports.
Partner summaries
Each partner gets its own retained-variant summary, including panel genes, secondary findings, and X-chromosome review where relevant.
Same-gene warnings
The couple-level table highlights genes where both partners still have retained heterozygous variants after filtering.
Exports
The workflow writes downloadable TSV, HTML, and Word outputs with separate findings for each partner and the couple summary.
The trio workflow extends proband prioritization with parental allele-state evidence. It is designed for family-aware interpretation: the proband is ranked first, and parental calls are then used to evaluate inheritance hypotheses. Missing parent files, low-quality parental calls, or absent coverage are treated as missing evidence rather than definitive evidence of absence.
Family intake. VarSage records the proband case details, phenotype information, genome build, sample IDs, expected inheritance, and the mother and father variant files when provided.
Proband-first prioritization. The proband undergoes the rare-disease workflow first, including phenotype matching, annotation, filtration, ranking, recessive-pair detection, provisional ACMG evidence, and optional AI review.
Parent file normalization. Mother and father files are parsed with the same coordinate and annotation logic, so parental calls can be compared against proband candidates using chromosome, position, reference, alternate, and gene context.
Inheritance-state comparison. For each important proband candidate, VarSage records whether the same allele is observed in mother, father, both parents, neither parent, or cannot be assessed because parental evidence is unavailable or inadequate.
Possible de novo review. Variants seen in the proband but not confidently seen in either parent are flagged as possible de novo findings. Parentage, locus coverage, genotype quality, sample identity, mosaicism, and orthogonal confirmation remain explicit review requirements.
Compound-heterozygous review. For recessive genes, VarSage looks for pairs of proband variants in the same gene that may have different parental origins. These pairs are summarized so the reviewer can inspect phase, partner-allele classification, genotype quality, and gene-disease fit.
Causative shortlist. The pipeline combines proband ranking and parent-aware evidence to identify the most compelling candidates for follow-up. A variant can carry more than one review label, such as causative and possible de novo, without being printed repeatedly.
Parent review tables. Each parent receives one concise retained-variant table. Category labels explain whether a variant is trio-relevant, phenotype-related, part of the carrier panel, a secondary finding, or part of female X-chromosome review.
Outputs. VarSage writes concise HTML and Word trio reports with key proband findings, compound-het pair summaries, parent review tables, warnings, methodology, and limits.
Proband shortlist
The proband still gets a full rare-disease ranking table with phenotype, inheritance, and evidence context.
Parent-aware findings
The trio summary separates possible de novo, compound-heterozygous, and causative candidate views from the parental comparisons.
Family reports
Mother and father each receive their own structured summary, and the final exports include the trio-specific interpretation sections.
Every gene-bearing result view in the web application uses the same on-demand GTEx panel. Rare disease ranked candidates, Carrier screening same-gene warnings and partner tables, and Trio proband, inheritance-finding, compound-heterozygous, causative, parent-review, and literature-context tables all expose the same Expression control.
Opening the control retrieves the gene's GTEx Portal profile and shows median TPM by tissue, tissue ranking, expression breadth, tissue specificity, a distribution summary, phenotype-linked tissue context when a broad local mapping is available, a locally saved reviewer note, and a link to the GTEx gene page. Expression data are presented as biological plausibility context rather than as variant-classification evidence.
Data source
VarSage uses GTEx Portal API v2, with gtex_v8 and GENCODE v26 as the defaults. Expression is loaded only when a signed-in user opens a gene panel, and unavailable remote data does not interrupt the analysis.
Privacy boundary
The authenticated VarSage endpoint receives the selected gene and phenotype labels. VarSage sends the gene lookup to GTEx and performs phenotype-to-tissue matching locally; patient files, patient IDs, and clinical narrative are not sent to GTEx by this feature.
Interpretation boundary
Expression is biological context, not an ACMG/AMP criterion and not a pathogenicity measure. It does not change filtering, ranking, classification, or reportability. Low bulk-tissue expression does not exclude a gene because cell type, developmental stage, transcript, and disease state matter.
Why it is shown separately: GTEx expression can support a reviewer’s biological plausibility assessment, but it cannot substitute for gene-disease validity, transcript relevance, mechanism, functional evidence, segregation, phase, or laboratory confirmation.
VarSage uses the Human Phenotype Ontology (HPO) as the controlled vocabulary for clinical features. HPO terms are used for rare-disease ranking, phenotype-concern scope in carrier screening, trio proband interpretation, and gene-disease plausibility review. The system preserves the difference between observed findings, explicitly absent findings, and inferred ontology matching context.
How HPO terms are collected
The matcher recognizes HPO IDs, canonical names, synonyms, selected clinically equivalent aliases, punctuation-normalized terms, and common plural forms. HPO hierarchy is used when mapping phenotypes to genes: for example, Migraine (HP:0002076) is a child of Headache (HP:0002315), so a gene annotated with Headache can remain in scope for a patient with migraine. The extractor may add direct, clinically useful parent terms such as Headache to the editable suggestion list; users can remove any suggestion before submitting the case. Co-symptoms such as photophobia, nausea, visual impairment, or aura are not inferred unless the source text states them or the user confirms them. AI suggestions still require a valid ontology ID and supporting text, and explicit negation takes precedence.
Manual entry
Type HPO terms directly into the Known HPO Terms field. The autocomplete searches by term name and synonym. Add terms one at a time; they appear as removable chips below the input.
AI extraction
Use Extract HPO terms with AI below each pipeline's clinical description. The model suggests present and explicitly absent terms, including strongly supported ontology parents, while deterministic validation requires valid identifiers and evidence from the submitted text. Suggestions appear in removable chips and remain editable before the analysis starts.
Absent (excluded) symptoms
Terms entered in the Absent HPO Terms field are recorded as excluded findings. They help refine the phenotype profile but do not contribute positively to the ranking score.
Text matching
The deterministic matcher normalizes text, handles common plurals (seizures matches Seizure), and searches HPO term names and synonyms. It prioritizes more specific (longer) terms when the 25-term cap is reached. Gene-scope matching additionally recognizes a patient term's HPO ancestors, but it never treats unrelated associated symptoms as observed findings.
Gene-phenotype associations
VarSage loads a gene-phenotype mapping file configured through GENE_PHENOTYPES_FILE. HGNC aliases and HPOA disease annotations can enrich this layer by connecting previous gene symbols, OMIM/ORPHA disease identifiers, disease-level phenotypes, and known inheritance terms. Phenotype overlap is therefore a structured plausibility measure, not a proof of causality.
Standard Variant Call Format. VarSage reads genotype, depth, genotype quality, filter status, INFO annotations, FORMAT fields, and transcript annotations. Multi-sample VCFs are supported when the analyzed sample ID is selected.
TSV / CSV
Tabular files with separate coordinate columns or a single parseable variant column. Common aliases for chromosome, position, reference allele, alternate allele, gene, consequence, and frequency are recognized.
TXT
Plain text files are treated as delimited tables when a consistent separator is present. They are most appropriate for small, already-curated variant lists.
Auto-detection
VarSage can infer the parser from filename and content. For non-standard exports, optional AI schema detection can identify columns, after which coordinate parsing remains deterministic.
BED region filtering
BED files define genomic intervals of interest, usually assay targets or capture regions. When a BED target set is selected, variants outside those intervals are removed before Jannovar annotation and external database merging. This improves runtime and keeps the report aligned with the technical scope of the assay.
BED coordinates use the standard 0-based half-open convention, while VCF and most variant tables use 1-based variant positions. VarSage handles the comparison internally, but the BED file must correspond to the same genome build as the uploaded variants.
Note: BED filtering is build-aware. A BED file configured for hg38 is only offered when the run uses hg38. The same catalog is shared across rare-disease, carrier-screening, and trio workflows.
The VarSage SuperDB layer is a local, build-specific coordinate annotation resource. It is intended to provide high-value fields needed for rapid interpretation: population rarity, clinical assertion context, prediction scores, splice scores, and stable variant identifiers. Large SuperDB files can be BGZIP-compressed and Tabix-indexed so only positions present in the current analysis are queried.
What super databases provide
maxFreq
maxFreq summarizes population rarity by retaining the maximum observed allele frequency across configured population resources for the same allele. It provides a conservative rarity screen, but disease prevalence, penetrance, ancestry, and inheritance mode still determine the appropriate clinical threshold.
Prediction scores
REVEL, CADD, SpliceAI, InterVar-like fields, ClinVar-like assertions, and custom laboratory columns can be attached when present in local annotation tables. These fields support review but do not replace primary evidence assessment.
VarSage offers optional AI-assisted features, but the platform remains functional without any model call. AI can help translate clinical text into reviewable HPO candidates, identify unusual table schemas, summarize evidence for top-ranked variants, and draft cautious report narrative from structured data. It is not the authority for coordinate extraction, allele identity, ACMG evidence, or clinical classification.
What AI can do
Phenotype extraction
AI reads the free-text clinical description and suggests present or explicitly absent HPO terms. Suggestions must resolve to valid ontology identifiers and should be interpreted as phenotype curation assistance.
Evidence review
When enabled for a case, AI reviews a bounded top-ranked candidate set plus selected rescue candidates and provides a narrative summary of phenotype fit, inheritance support, evidence quality, and missing confirmation steps.
ACMG/AMP refinement
For AI-reviewed candidates, the model may suggest ACMG/AMP criteria that appear compatible with the supplied structured evidence. Deterministic validation checks the code, strength, and prerequisites before accepting any provisional refinement. Expert verification remains mandatory.
Schema detection
For unusual TSV/CSV headers, AI can identify which columns appear to encode chromosome, position, reference, alternate, gene, consequence, quality, and annotation fields. VarSage then parses coordinates deterministically.
Case conclusion
When enabled, AI drafts a case-level synthesis from structured candidate data. The synthesis separates review priority, provisional classification, case-level causality, reportability, and missing evidence.
Important: AI output is decision support. A model cannot establish pathogenicity, create ACMG evidence by assertion, override deterministic allele matching, or make a variant reportable without explicit supporting evidence and expert review.
VarSage provides conservative ACMG/AMP-oriented evidence summaries for germline Mendelian sequence variants. These summaries are provisional decision-support aids, not clinical laboratory classifications. Evidence is deliberately labeled by state so reviewers can distinguish explicit auditable data from hypotheses and contextual information.
Evidence categories
Phenotype evidence
Derived from HPO overlap between the patient profile and gene-disease associations. High overlap supports phenotype-gene plausibility, but it is not sufficient by itself for variant pathogenicity.
Inheritance evidence
Based on zygosity, sex, expected inheritance mode, known gene inheritance, and family clues. Trio and phase evidence remain provisional until sample identity, parentage, callability, and segregation are reviewed.
Population evidence
Derived from maxFreq and other frequency fields. Rarity supports review, but maximum credible allele frequency depends on prevalence, penetrance, inheritance mode, and disease-specific guidance.
Prediction evidence
REVEL, CADD, SpliceAI, and similar predictors support computational-impact review only when their thresholds and applicability are appropriate. Correlated predictors should not be counted independently.
Quality evidence
Depth, genotype quality, and filter status are used to reduce technical noise. Quality metrics support confidence in a call, not pathogenicity.
Evidence states
State
Meaning
Example
Applied
The prerequisite fact is explicitly present in the source data and can be audited.
A supplied validated functional assay strength, or de novo evidence with parentage confirmation.
Candidate
The data suggest a criterion, but disease-specific or laboratory prerequisites are incomplete.
Rarity compatible with PM2_Supporting, or possible compound heterozygosity without confirmed trans phase.
Informational
The field is useful context but is not counted as ACMG/AMP evidence.
ClinVar label, GTEx expression, LitVar context, or AI narrative without primary supporting evidence.
How to interpret ACMG evidence
PVS1 is not assigned from consequence alone; loss-of-function disease mechanism, transcript relevance, NMD expectation, exon rescue, and the ClinGen PVS1 decision tree require review. PM2 is treated conservatively at Supporting strength. PP3 and BP4 require calibrated predictors. PP5 and BP6 are not generated from ClinVar labels. De novo, trans-phase, co-segregation, and functional evidence require the corresponding laboratory and disease-specific prerequisites.