Skip to content

TrialMatchAI

Patient profiles, retrieval, criterion evidence, and ranked reports

A configurable CLI for patient-to-trial retrieval and eligibility review. TrialMatchAI imports patient information, searches local LanceDB tables, and writes ranked results and HTML reports. The default model workflow includes entity extraction, reranking, and criterion-level eligibility assessment. Query expansion is separately configurable.

Beta release — declared validation scope

The release gates test the installed CLI with synthetic data and real CPU search. GPU inference, clinical accuracy, and production clinical deployment remain separate qualification work. Review generated assessments against source evidence; a retrieval score is not an eligibility determination.

Start with the synthetic demo

Install with Python 3.11 in an activated virtual environment:

python -m pip install trialmatchai==0.9.1
trialmatchai demo --workdir ./demo-workspace
trialmatchai demo --workdir ./demo-workspace --resume

Open the report path printed by the command. This walkthrough uses BM25 and hashing embeddings with model stages disabled. It creates synthetic records only; resume preserves completed ranking and repairs a missing or truncated patient report.

For your own data, use the README setup instructions. Model stages require additional dependencies and configuration. The default adapter IDs currently need replacement with downloaded local paths. Format mappings, cache invalidation, and registry freshness have documented limits.

Guides

Guide What it covers
Architecture Runtime components and local storage
Pipeline and CLI Stages, presets, and report commands
Patient interoperability Supported formats and mappings
Registry updater Fetching and indexing registry studies
Fine-tuning Training entry points and data formats
Paper results and reproduction Ranking results, pooled recall, ideal candidates, and downloadable values
API reference Python interfaces
Release runbook Checksums, bootstrap recovery, CI, and publishing
Validation record Checks performed and their limits
Codebase review Historical audit evidence
Production roadmap Sequenced clinical, ranking, agent, and deployment work

The current pipeline follows a fixed stage sequence. Bounded agents and a clinical review workspace are planned. This release does not claim newly improved benchmark scores; the paper describes the published research, while the validation record describes this software release.