Tools

Software we build and ship.

Open-source scientific software from our lab, built to be used, audited, and extended by other researchers.

Biomedical tools

From variants and cells to molecules.

Open Source
Python package · scRNA-seq

msep

Multi-scale entropy profiling for single-cell transcriptomics.
pip install msep

Multi-Scale Entropy Profiling integrates per-cell Shannon entropy with across-cells coefficient of variation (CV), decomposed by biological pathway, to characterise population-level transcriptomic coordination in single-cell RNA-seq data. The framework reveals states invisible to single-scale analyses — such as populations that are individually diverse (high per-cell entropy) yet collectively disciplined (low across-cells CV) — and is applicable to any cancer type or cellular system.

  • Pathway-decomposed entropy, not a single global score
  • 33,000+ curated gene sets via MSigDB out of the box
  • Bootstrap confidence intervals reported for every estimate
  • Built for standard scRNA-seq workflows
Open Source
Clinical decision support

VUS Lens

Ancestry-aware variant interpretation, built to expose the gaps.

VUS Lens is a deterministic ACMG classification engine that does something most variant interpretation tools do not: it audits whether the evidence behind a call is actually strong enough for the patient's ancestry, rather than assuming Western-centric reference databases generalize across populations. Its reasoning layer is built with Claude.

  • Deterministic ACMG rule engine — same input, same output, every time
  • Ancestry-confidence auditing surfaces where reference data falls short
  • Validated against 1,277 known-pathogenic variants with zero false-benign calls
  • Reasoning layer built with Claude
On request
Clinical decision support

VUS Pipeline

Transparent, ACMG-compliant variant classification for the clinic.

VUS Pipeline is a production-grade clinical decision-support system that applies the ACMG/AMP 2015 + ClinGen SVI + Tavtigian point framework deterministically, gathers evidence in parallel from authoritative sources, and layers advisory mechanistic interpretation on top — every step traceable and source-cited. The class is set by rules, never by an LLM.

  • Deterministic ACMG / SVI decision engine (Tavtigian points)
  • 8-layer parallel evidence collection (gnomAD v4, ClinVar, SpliceAI, ClinGen, CIViC)
  • Advisory mechanistic interpretation layer — never changes the class
  • Source-cited reports (HTML / PDF / DOCX) with a citation-verified assistant
Open Source
Python · Tox21 · cheminformatics

HybridTox

Safe-by-Design toxicity prediction that fails safely on unseen scaffolds.

HybridTox integrates the explicit structural memory of Random Forests with the topological intuition of Graph Neural Networks through a transparent stacking ensemble that acts as a Mixture of Experts. It is built to prevent deep learning's silent failures — high-confidence errors on out-of-distribution chemistry — and to act as a fail-safe in data-scarce regimes.

  • Random Forest + GNN experts, combined by a transparent logistic-regression stacker
  • Evaluated on 12 Tox21 endpoints with a 5-seed protocol
  • Nested scaffold split — tested on chemistry the model has not seen
  • Preprint on ChemRxiv; under review at the Journal of Chemical Information and Modeling
Open Source
scRNA-seq + LLM agents · chordoma

Dual Shield

A two-layer model of immune evasion and structural resistance in chordoma.

Dual Shield is a single-cell RNA-seq and LLM-agent pipeline for chordoma. It describes an outer immune-evasion shield (HLA-E/NKG2A) and an inner structural-resistance shield (vimentin-driven partial EMT), and searches for combination strategies that could dismantle both. Agent-generated hypotheses pass through an adversarial review before they are reported.

  • Single-cell evidence for both shields from chordoma scRNA-seq
  • LLM-agent hypothesis generation with adversarial review
  • Candidate combinations designated for 3D organoid co-culture validation
  • Apache-2.0 code; restricted compound-level outputs kept private

Research hypotheses only — not clinical guidance.

Auditable agent systems

Agents whose reasoning you can inspect.

The same principle behind our clinical tools — deterministic checks decide, models assist, every step is logged — applied to agentic AI. These are research prototypes, not clinical products.

Multi-agent evidence monitoring · Google ADK

Recall

A deterministic controller and policy gate hold all authority; an Evidence Watcher, an Evidence Assessor and a Citation Auditor stay narrowly scoped, so discovery, interpretation and citation checking cannot share failure modes. Only the policy gate can return NO_ACTION, ABSTAIN or REVIEW_REQUIRED.

Non-clinical research prototype: it never classifies patients, edits reports or contacts patients.

Professional agent · Agents for Humans Hackathon

Watershed Memory

One watershed case carries observations, operator reviews and approved field plans forward from storm to storm, and missing station evidence gets its own review. Every step leaves a tool receipt you can inspect.

The demo uses saved USGS readings and simulated field work; no sign-in required.

ML lineage agent · DataHub

datahub-ml-guard

Writes ML lineage from training code into DataHub, then walks the column-derivation graph to judge suspicious feature–label relationships. Findings are written back into the graph as evidence, together with the parts of the estate the agent could not see.

Live instance requires sign-in. Upstream contribution to DataHub (PR #18993).