From target binding to neuropharmacological effects: what should count as a credible computational forecast?
One of the central problems in computational neuropharmacology is deciding what we actually mean when we say that a model can “predict” or “forecast” the effects of a compound on the brain.
Target-binding data are an obvious starting point, but they are rarely the endpoint. Two compounds acting at the same nominal receptor can produce substantially different biological outcomes because affinity is only one variable in a much larger system: intrinsic efficacy and functional selectivity, receptor and transporter distribution, cell type, downstream signaling, pharmacokinetics, off-target activity, network state, and interactions with other signaling systems can all affect the eventual phenotype.
The difficult scientific problem stretches well beyond the question of:
What does this molecule bind to?
A more comprehensive investigation must entail:
How much of the causal path from molecular interaction to cellular response, circuit-level change, and ultimately cognition or behavior can we predict before observing the outcome?
That distinction seems increasingly important as biological knowledge graphs, foundation models, graph neural networks, molecular models, and multimodal neuroscience datasets are combined into systems that attempt increasingly ambitious biological inference.
A possible standard: prospective, multiscale prediction
A convincing benchmark should probably require more than recovering facts that are already represented somewhere in the training data or underlying knowledge base.
Imagine giving a system a set of held-out compounds and requiring it to make predictions at several biological scales before the relevant experimental observations are revealed.
At the molecular level, the system might predict target engagement and direction of modulation. At the cellular level, it might predict affected pathways or perturbational signatures. At the anatomical level, it might identify the brain regions or cell populations most likely to be affected. At the systems level, it might predict changes in functional networks, electrophysiology, or other measurable neural states. Finally, at the phenotypic level, it could forecast cognitive, behavioral, therapeutic, or adverse effects.
The important feature is that these predictions would be committed in advance, together with confidence estimates.
This would let us evaluate not only whether a system produces biologically coherent explanations, but whether those explanations actually constrain future observations.
Why multiscale validation matters
There are already important pieces of the necessary infrastructure.
Large perturbational resources such as the Connectivity Map demonstrate that molecular interventions can be organized by downstream biological response rather than chemical structure alone. Human neuroimaging resources increasingly make it possible to relate receptor and transporter distributions to large-scale brain organization. Multimodal cortical parcellations provide increasingly precise anatomical coordinate systems, while tools such as neuromaps make it easier to compare molecular, structural, and functional brain maps across modalities.
What is less clear is how these scales should be connected when evaluating a forecasting system.
A prediction could be “correct” about a molecular target but wrong about the resulting network effect. It could correctly identify a brain region while predicting the wrong direction of functional change. It could generate a plausible mechanistic chain yet fail completely at the behavioral level.
Those failures are scientifically informative. A useful benchmark should preserve them rather than collapsing everything into a single accuracy number.
It should also evaluate calibration. A system that knows when the available evidence is weak is scientifically more useful than one that produces equally confident answers for well-supported and poorly constrained predictions.
The data-leakage problem
There is another complication: retrospective biological prediction is unusually vulnerable to information leakage.
For established compounds, target profiles, pathway annotations, transcriptomic signatures, imaging results, clinical effects, adverse events, and even mechanistic interpretations may all ultimately derive from overlapping literature.
A sufficiently capable model may reconstruct the answer without actually performing the kind of cross-scale inference we think we are testing.
For that reason, the strongest evaluation may ultimately have to be prospective: predictions are timestamped before a new experiment, dataset, or compound characterization becomes available, and evaluated only afterward.
That is considerably harder—but it would turn biological forecasting into something genuinely falsifiable.
Questions for the OpenLabs community
I would be particularly interested in views on four questions:
- What should the ground truth be? Should a CNS forecasting benchmark prioritize molecular assays, perturbational transcriptomics, PET/fMRI/EEG, behavioral phenotypes, clinical outcomes, or some explicitly multiscale combination?
- How should partially correct mechanistic predictions be scored? If a model identifies the correct receptor and brain system but predicts the wrong downstream phenotype, that is different from being wrong at every scale.
- How can we distinguish mechanistic generalization from sophisticated retrieval? Is rigorous temporal holdout sufficient, or do we ultimately need prospective experiments on genuinely new interventions?
- What should uncertainty look like? Should forecasts provide calibrated probabilities for individual claims, confidence intervals over quantitative effects, explicit alternative mechanisms, or all three?
Our interest at Nootropics DAO comes from working on computational approaches to reasoning across compounds, receptors, signaling pathways, brain regions, and cognitive effects. But the broader methodological question seems more important than any particular architecture.
A system should not be considered scientifically credible simply because it can produce a coherent mechanistic story.
It should earn credibility by making predictions that could have been wrong.
Selected references
Hansen JY, Shafiei G, Markello RD, et al. Mapping neurotransmitter systems to the structural and functional organization of the human neocortex. Nature Neuroscience. 2022. DOI: 10.1038/s41593-022-01186-3.
Markello RD, Hansen JY, Liu ZQ, et al. neuromaps: structural and functional interpretation of brain maps. Nature Methods. 2022;19:1472–1479. DOI: 10.1038/s41592-022-01625-w.
Glasser MF, Coalson TS, Robinson EC, et al. A multi-modal parcellation of human cerebral cortex. Nature. 2016;536:171–178. DOI: 10.1038/nature18933.
Subramanian A, Narayan R, Corsello SM, et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. Cell. 2017;171:1437–1452.e17. DOI: 10.1016/j.cell.2017.10.049.