Understanding the question
An AI-selected biomarker becomes useful only when its measurement and interpretation are supported for a defined purpose. Candidate discovery can reveal patterns associated with a condition or outcome, but association alone does not establish a reliable test. Validation must address analytical performance, relevance to the intended population and whether the information improves the decision it is meant to support. A model can perform well in a curated research dataset and fail in a different setting because of selection effects or technical variation. Independent evaluation and a clear intended use are therefore essential before making diagnostic or prognostic claims.
What a useful investigation needs to consider
Define whether the candidate is intended for diagnosis, prognosis, monitoring or another research purpose. These uses require different populations, endpoints and timing. Avoid selecting features using the entire dataset before reporting performance on a supposedly independent subset.
Examine calibration, subgroup behavior and performance at relevant prevalence, not only discrimination metrics. Technical batch effects and treatment history can masquerade as biomarkers. Separate research findings from validated clinical tests, and ensure that any future clinical claim is supported by appropriate evidence and regulatory assessment.
Read the detailed explanation
The companion article explores ai biomarker discovery: designing validation before selection in more depth, with topic-specific explanations and source material.
AI biomarker discovery: designing validation before selectionSources and further reading
These resources provide background and methods relevant to this topic. They are not evidence of a FormBio product or a personalized recommendation.