How can we avoid Goodhart's Law in DeSci?
Below is an article I wrote about Goodhart's law, which put simply is: when a measure becomes a target, it ceases being a good measure.
On one hand, metrics fuel science analysis. On the other, optimizing for numbers can make us miss the forest for the trees.
How can we optimize for what actually matters in DeSci?
Full article below:
Somewhere between breakthrough and acceptance, many promising ideas quietly fade away. They failed not because the treatment did not work.
They simply did not fill the checklist.
☐ Patentable
☐ Treats a well defined condition
☐ Significantly better than existing treatments
And at some point you start to wonder: when did the process become more important than the purpose? When did the rules become more sacred than the lives we’re trying to improve?
Goodhart’s Law states that
“When a measure becomes a target, it stops being a good measure.”
In other words, if we focus too much on hitting certain numbers or metrics, we risk losing sight of what really matters. In research, this can slow progress, block innovation, and even keep effective treatments from reaching those who need them.
When Rules Take Over Imagine you’ve found a therapy that helps people. You’ve seen the improvements with your own eyes. Patients talk about improved quality of life, restored function, and relief.
But then the rules arrive:
- “Your sample size wasn’t large enough.”
- “You didn’t blind this in the standard way.
- “This doesn’t fit the conventional model of intervention.”
- “Your outcomes don’t map to regulatory endpoints.”
The therapy isn’t judged on whether it helps people, but whether it fits the framework for proving it helps people. Real progress in research is often messy, complex, and full of surprises, and yet, rigid adherence to metrics often decides which ideas move forward and which are left behind.
Worse yet: if a small group of people control what is considered a legitimate trial, then it becomes practical for corporations to capture them for their own benefit.
The Story of MAPS and MDMA Therapy
A clear example comes from the Multidisciplinary Association for Psychedelic Studies (MAPS) and its research on MDMA-assisted therapy for PTSD. For years, patients who didn't respond to traditional mental health treatments found relief through MDMA therapy combined with guided sessions. The results were powerful and obvious to clinicians and patients alike. The therapy worked in practice: people got better, and therapists noticed.
Yet approval ultimately failed, not because the therapy didn't work, but because it couldn't satisfy a methodological standard designed for conventional pharmaceuticals. The FDA's gold standard for drug approval requires "adequate and well-controlled" studies, which traditionally means randomized, double-blind, placebo-controlled trials where neither patients nor researchers know who received the real drug. This works for most medications - a blood pressure pill feels the same as a sugar pill. But MDMA is impossible to blind in any meaningful way.
MAPS ran two Phase 3 trials (MAPP1 and MAPP2) that the FDA initially allowed to proceed under a double-blind, placebo-controlled design. The agency even reviewed and approved the protocol. But from the beginning, FDA reviewers repeatedly warned about "functional unblinding" and never committed to the trial counting as sufficient evidence. They suggested using an active comparator, something like niacin or low-dose MDMA that would produce noticeable effects, to better preserve blinding. MAPS argued these alternatives might worsen PTSD symptoms, and the two sides never agreed on a solution.
When the trials were complete, the blinding survey revealed what everyone already knew was inevitable: 94% of participants who received MDMA correctly guessed they'd gotten the real drug. 75% of placebo recipients knew they'd gotten placebo[1]. The therapists conducting the sessions could obviously tell too, as MDMA produces hours of unmistakable psychological and physiological effects. Many participants had previous recreational experience with MDMA, making recognition even easier.
In June 2024, an FDA advisory committee voted 9-2 that the trials had not demonstrated substantial evidence of effectiveness, and 10-1 that benefits did not outweigh risks. The committee's reasoning was clear: when virtually everyone knows what they received, expectation bias becomes impossible to rule out. Perhaps people improved because they believed they were getting a powerful treatment, not because MDMA itself was therapeutic. Two months later, the FDA issued a Complete Response Letter formally declining approval. In September 2025, the agency publicly released detailed documentation explaining that the failed blinding undermined the entire evidential basis, even though the efficacy results were statistically positive.
Here's where Goodhart's Law reveals itself. The measure—perfect blinding in clinical trials—was designed to eliminate bias and ensure we can trust that drugs really work. It's a good measure for most drugs. But the FDA made this measure into an absolute requirement, an inflexible target that had to be met regardless of context. MDMA became a victim of a rule meant to protect patients. The drug's obvious psychoactive effects made the traditional measure impossible to satisfy, yet the agency insisted on it anyway. The system prioritized procedural purity over the practical question of whether the therapy actually helps people suffering from severe PTSD.
When a measure becomes a target, it stops being a good measure. The clinical trial blinding requirement exists to reveal truth, but when applied rigidly to a drug that cannot be blinded, it obscures truth instead. MAPS found itself in a regulatory catch-22: they couldn't prove MDMA worked without perfect blinding, but perfect blinding was pharmacologically impossible. The measure had become the obstacle, and human benefit -the whole point of drug development - was lost in the process.
Why Goodhart’s Law Matters to All of Us Goodhart’s Law isn’t just about numbers or regulations; it’s more about how we define progress. When science focuses on meeting the “right numbers” instead of real results, people and patients are left behind. The case of the MDMA study was emotionally illustrative, but let us turn now to most abstract ways in which this principle impacts longevity research, and knowledge creation more broadly.
Biomarkers of Aging: The Seductive Proxy
Biological age clocks, such as epigenetic methylation patterns, telomere length, inflammatory markers, and metabolomic signatures promise something irresistible: a number that tells you how old you really are, independent of chronological time. More seductively, they promise a shortcut: instead of waiting decades to see if an intervention extends lifespan, one could just measure whether it "reverses biological age" and you'll know in months.
This is both genuinely useful and profoundly dangerous.
The utility is real. Lifespan studies in dogs take 10-15 years. In humans, they take longer than most research careers. If you're testing follistatin gene therapy for canine arthritis, you cannot wait until 2040 to know if treated dogs lived longer. You need intermediate signals like:
- Does the dog move better?
- Do inflammatory markers improve?
- Does the epigenetic clock slow?
These are valuable, actionable clues that let you iterate and learn.The danger emerges when clues become destinations. And suddenly we are back to asking; are we actually slowing aging or just moving numbers?
The fundamental problem is we don't know which biomarkers of aging track actual aging versus which ones just track things correlated with aging in the populations where they were measured. Epigenetic clocks are built on associations. While this methylation pattern is common in 70-year-olds, and this one in 30-year-olds, there is still only an association, and not a mechanism. If you change the methylation pattern without changing the underlying biology, have you reversed aging, or just gamed the metric?
Some interventions might genuinely slow aging and show it in biomarkers. Others might alter biomarkers through different mechanisms entirely: acute stress responses, metabolic shifts, changes in cell composition. The biomarker moves, but the dog doesn't live longer or suffer less. You've optimized the proxy, not the outcome.
The stakes are especially high because biological age has become a funding and legitimacy gate. Longevity startups pitch investors on "we reduced biological age by X years in Y weeks." Clinics sell aging biomarker panels as diagnostic tools. Supplement companies market products around clock-reversal claims. The incentive structure now rewards interventions that move biomarkers dramatically and quickly, regardless of whether they produce durable health benefits.
This doesn't mean aging biomarkers are worthless. On the contrary: the field of biomarkers of aging will eventually converge on the best indicators of vitality we've ever found. Longitudinal studies that track both biomarkers and actual outcomes, disease onset, functional decline, lifespan will reveal which markers are faithful proxies and which are mirages.
The path forward requires epistemic humility. Use aging biomarkers as one signal among many, but never as the sole measure of success. Pairing them with functional outcomes like:
- Can the dog still run?
- Does the human maintain muscle mass and cognitive sharpness?
- Are they free from chronic pain?
Track biomarkers longitudinally to see if changes persist or revert. Most critically, always circle back to the ground truth: healthspan and lifespan.
For Dog Years, this means treating biological age clocks as useful exploratory tools, but not as validation endpoints. If follistatin therapy improves mobility, reduces inflammation, and extends active lifespan in treated dogs, that's success, even if some aging biomarker moves in an unexpected direction. Conversely, if a biomarker shifts dramatically but dogs don't actually live better or longer, you haven't succeeded; you've just found an intervention that games that particular metric.
The field will mature. Ten years from now, with enough longitudinal data across diverse populations and interventions, we'll know which biomarkers are robust and which were training-set artifacts. We'll have validated clocks that genuinely track the pace of aging, not just its correlates. We'll understand the mechanism well enough to distinguish "this intervention slows aging" from "this intervention acutely stresses cells in ways that shift methylation patterns."
Until then, biological age remains a seductive proxy: useful as a clue, dangerous as a target. The mission is more good years for dogs and humans. Biomarkers are tools in service of that mission, valuable when they illuminate the path, misleading when mistaken for the destination.
Don't let the clock become what you're trying to stop.
Other areas affected by Goodhart’s Law:
1. AI Alignment
In AI safety, most companies now optimize to pass safety tests, not necessarily to actually be safe. And sometimes it feels like machines are trained more to look harmless than to be harmless. As AI capabilities scale and regulatory frameworks emerge, we're at risk of building entire governance systems around metrics that labs have already learned to game. Safety scores become theater which is reassuring to policymakers, but meaningless in practice.
2. Citation Counts and the H-Index Economy
In the world of academics today, researchers chase citation counts like arcade points, sometimes citing each other in cozy loops, leaving us to wonder if the game is about knowledge or reputation.
The tragedy is that this crowds out risky ideas, negative results, replication studies, and slows the accumulation of meaningful data.
3. “Decentralization" as Virtue Signal
In crypto and decentralized governance, there’s a tendency to perform decentralization, as if broadcasting the appearance of community control matters more than ensuring real participatory decision-making.
The clear problem is that decentralization is not intrinsically good on its own. It's a tool that trades efficiency and speed for resilience, censorship-resistance, and stakeholder alignment. Sometimes that trade is worth it. Often it's not. A decentralized clinical trial might be more resilient to institutional capture but slower to adapt protocols when safety signals emerge.
The question should never be "how decentralized is this?" but rather "what problems does decentralization solve here, and what does it cost?" When DeSci projects optimize for decentralization metrics, token distribution curves, number of governance participants, and voting frequency rather than asking whether decentralization serves the mission, Goodhart's Law has won.
For Dog Years specifically, this means: use DeSci structures where they genuinely help (fundraising, participant incentives, transparent data), but don't let "looking sufficiently DeSci" to crypto investors become the goal. If tri-cameral governance or token-weighted voting creates overhead without improving outcomes for arthritic dogs, you're optimizing for the wrong thing.
When the metric becomes the mission, the mission quietly dies.
4. IQ and the Galaxy-Brained Policy Machine
The IQ discourse online has a peculiar feature: it is rarely about the test itself. It is about the chain of inference hung from it. Take a score, average it across a population, attach a heritability estimate, and suddenly a single number is doing the work of an entire theory of society. Each link in that chain is weaker than the last, but by the end the conclusion is delivered with the confidence of arithmetic.
This is the same move we warned about with biological age clocks. A human mind, like a dog's aging body, is a sprawling, multidimensional thing. Compressing it into one number makes it tractable. Compression also throws information away, and what's left looks like precision because it has decimal places. The number stops being a summary of the thing and starts being treated as the thing.
Goodhart shows up almost immediately. Raw IQ scores rose roughly three points per decade across the 20th century, the Flynn effect. Almost no one concludes that their grandparents were borderline impaired, so the number clearly moves for reasons other than the thing people claim it measures. Schooling raises scores, and a 2018 meta-analysis put the gain at a few points per additional year of education. Test familiarity raises scores. Nutrition raises scores. The measure is responsive to exactly the environmental levers that the "it's all baked in" crowd says don't matter.
Then there are the national IQ datasets that circulate as infographics. When researchers audited some of the most widely cited estimates, they found that several countries' figures rested on small, unrepresentative, or decades-old samples. Those samples were then extrapolated to entire nations and fed into regressions about GDP. That is a colorful map built on a questionable foundation.
Heritability is the other load-bearing confusion. It is a population statistic describing variance within a given environment, not a statement about how fixed a trait is. Height is highly heritable, and average height still rose dramatically once nutrition improved. PKU is about as genetic as a condition gets, yet a diet prevents the intellectual disability it would otherwise cause. "Genetic" does not mean "immutable," and it certainly does not mean "therefore this policy."
Which brings us to Vitalik Buterin's test of galaxy brain resistance: how easily can a style of argument be bent to justify whatever you had already decided? The IQ-as-policy argument scores close to zero. It has been used to justify immigration restriction and immigration expansion, cutting school funding ("it's genetic anyway") and increasing it, eugenics programs and embryo-selection startups. When the same number can be recruited for any conclusion, the number isn't the reason for the conclusion. Rather, it's just a decoration. Revealingly, the policy preference almost always arrives first, and the psychometrics are deployed to justify it afterwards.
None of this requires pretending cognitive testing is useless. It requires noticing when a measure has been promoted from instrument to judge. We wouldn't let one methylation score decide whether a dog is thriving. We shouldn't let one test score decide what a person, or a population, deserves.
Don't let the score become the person.
If We Must Measure Then We Must Measure What Matters
Metrics aren’t the enemy. They are incredibly useful when applied thoughtfully and interpreted responsibly. They help us observe patterns, detect subtle biological signals, and refine hypotheses. But they should never outrank lived, functional outcomes.
“Did the biomarker shift?” is insignificant compared to “Is the dog experiencing better movement, energy, comfort, and engagement with life?”
A biomarker might suggest biological improvement, but real improvement is visible in behavior and function. Metrics should be instruments, not judges. A methylation score can reveal interesting biological directionality, but does not prove meaningful impact. So when we observe:
- increased mobility
- reduced signs of pain or stiffness
- renewed interest in activity
- better responsiveness
- more social/affectionate behavior
- sustained physical comfort
Those observations are not “soft” or secondary; they are biological outcomes expressed through behavior. And objective measurement matters but so does the real-world manifestation of health.
Community Engagement vs. Genuine Participation
This same metric dynamic can surface in decentralized communities as well. In many DAOs, activity metrics, message volume, voting turnout, or wallet participation, can unintentionally become proxies for engagement. While these indicators are useful, they do not always capture the underlying quality or relevance of contributions.
True community participation often looks quieter and more focused. It may be seen in:
- a veterinarian raising a thoughtful methodological question
- a researcher carefully validating a dataset
- a dog owner providing consistent health observations
- contributors working behind the scenes on analysis or documentation
These efforts don’t always produce high visible activity, but they meaningfully advance the mission.
For Dog Years DAO, this distinction matters. Our goal is not simply to have a busy community, but a purpose-aligned community — one where participants contribute according to their expertise, capacity, and genuine interest in canine health and longevity.
Rather than measuring engagement by quantity of interaction, we can emphasize:
- clarity of input
- relevance of contributions
- quality of data
- constructive collaboration
- sustained commitment
- positive-sum participation
This approach supports an environment where professionals, practitioners, and caring dog owners feel comfortable engaging deeply, without pressure to perform activity for its own sake.
The intent is not to diminish visible participation, but to recognize that thoughtful engagement and quiet expertise are equally valuable and often essential to progressing the research in a responsible and credible way.
A Universal Pattern, Framed Diplomatically
Across many fields, from AI to academic research to decentralized governance, we see that meaningful intentions can become distilled into simplified metrics. And when metrics are emphasized too strongly, there is a natural risk that they begin to shape behavior.
The path forward is not to discard measurement, but to anchor it to purpose. Metrics should be supportive references, not substitutes for actual impact. When measurement remains subordinate to mission, we preserve both integrity and progress.
Why This is Crucial
Because an overreliance on metrics can lead to:
- optimizing for biomarker movement rather than functional improvement
- neglecting individual variation in aging expression
- misinterpreting short-term biological noise as long-term trend
- prematurely dismissing interventions that show real-world benefit
Metrics can illuminate biological possibility, but lived outcomes confirm biological reality. If the dataset says aging should be accelerating but the dog is visibly more active and comfortable then we must reconcile the discrepancy thoughtfully rather than dismiss the observation.
The Heart of Good Science
Good science isn’t sterile, it is compassionate and curious. It isn’t afraid of messiness, rather: it expects it.
The heart of good science should ask:
- What improves life?
- What reduces suffering?
- What restores function?
- What brings vitality?
- What brings hope?
And for any research project whether therapy for PTSD or longevity for dogs, the central question must remain: Is life getting better?
Everything else is secondary.
Metrics will always matter, but they are not the mission.
Whether it’s MDMA therapy for PTSD or longevity research for dogs, the point is to help life flourish, function, and thrive.
A New Way Forward: Open and Collaborative Research Regulations are not bad or useless. We need safety checks. We need rigor. But the point of science is not to protect the process, but to protect the people.
We need approaches that:
- embrace real-world evidence
- combine subjective and objective assessments
- accept dimensional rather than binary results
- reward real benefit rather than perfect alignment with protocol
This is why Dog Years DAO is setting an example of how research can evolve. By organizing community-driven studies, we can test ideas safely and responsibly without being slowed down unduly by bureaucratic processes.
Anyone who has ever loved a dog knows the bittersweet truth as their time with us is heartbreakingly short. We see them age in fast-forward. One year they’re bounding across parks and before we know it, they’re slowing down, sleeping more, and wincing as their joints stiffen. We can learn directly from real-life data on how dogs live, eat, and age, and then turn those into meaningful research.
Research then becomes practical, transparent, and useful for everyone.