How to Tell If a Health App's Predictions Are Real
Health apps have started talking about the future. Your recovery is trending down. You are likely to sleep poorly tonight. Expect lower readiness tomorrow. Some of that is real forecasting. A lot of it is a description of the past with the tense changed.
There is one test that separates the two, and it takes about a minute to apply to any tool you already use.
The test: was the claim made before the outcome, and was it scored afterwards
A prediction has two halves. It has to be stated in advance, and it has to be checked afterwards. A tool that does the first without the second is not predicting, it is speculating in public. A tool that does neither is describing.
Most consumer health software describes. That is not a criticism, description is genuinely useful, and the industry got very good at it. But describing yesterday and forecasting tomorrow are different engineering problems, and the marketing language for both has converged on the same words.
Four questions to ask any tool that claims to forecast
Was the claim timestamped before the thing it predicted? Open the app in the morning and read what it says about the day ahead. Then look at it again that evening. If the morning statement is gone, or has quietly been replaced with an explanation of what happened, you were reading a description.
Does it ever tell you it was wrong? This is the fastest disqualifier. A system that surfaces only its correct calls is not being evaluated, it is being marketed. Look for an explicit record of misses. If you cannot find one anywhere in the product, assume the misses exist and are simply not shown.
Is the comparison your own history or a population? A tool saying your readiness is low relative to a general population is telling you where you sit among strangers. A tool saying it is low relative to your own last six weeks is telling you something changed. Only the second can support a forecast about you.
Can you see what the forecast was based on? A prediction with no visible inputs cannot be argued with. You do not need the model weights, but you should be able to see which signals moved and in which direction.
Why retrospective explanations feel like predictions
The reason this is hard to spot is that hindsight is genuinely informative. When an app tells you at 8pm that your elevated resting heart rate explains your low recovery score, the explanation is probably correct. It reads as insight because it is insight. It just is not a forecast, and it carries no risk for the tool that produced it.
The tell is that a retrospective explanation is never wrong. It is fitted to an outcome that already happened. Anything that cannot be wrong also cannot be trusted in advance, because it has never been tested in advance.
| What the app shows you | What it actually is |
|---|---|
| Your recovery was low because your resting heart rate was elevated | A description, correct and unfalsifiable |
| Your recovery is likely to be low tomorrow, with no follow up | A speculation, never scored |
| Your recovery is likely to be low tomorrow, and here is yesterday's call marked as held or missed | A forecast under evaluation |
| A weekly summary of everything that happened | A record, which is a different and legitimate product |
What a scored forecast looks like in practice
Scoring only works if the possible outcomes are fixed in advance and include failure. The scheme we hold ourselves to has three states and nothing else: Held, Still open, or Missed. A call that came true is held. A call whose window has not closed yet is still open. A call that did not come true is missed, published as such, in the open.
The reason to prefer three states over a confidence percentage is that a percentage is unfalsifiable in a single instance. If a tool says sixty percent and the thing does not happen, it can always claim the forty percent branch. Held, still open and missed cannot be argued with after the fact, which is the point.
That commitment is why NuVARD is built to read across sources rather than score one in isolation: 300+ devices and apps feeding 53 physiological variables across 15 body systems, each with its own baseline and trend. A forecast about you is only checkable if the thing it is measured against is also about you.
What to do with an app that fails the test
Keep using it, and downgrade what you ask of it. A good recorder is worth having. The mistake is treating a description as a forecast and making decisions on that basis, which is how people end up rearranging a training week around a number that was never a claim about the future in the first place.
Three practical habits:
Read direction over days rather than the absolute number today. One morning is weather, a multi day drift is climate.
Look for corroboration across independent signals before you act. One metric moving is a question. Three moving together is an answer.
Write down what you expect before you look. If you are going to make a prediction from your own data, make it falsifiable in the same way you would demand of a tool.
Frequently asked questions
Can wearables actually predict anything
Some signals do carry forward looking information, particularly when several move together against a stable personal baseline. The open question is rarely whether prediction is possible in principle. It is whether a given product states its predictions in advance and reports its accuracy, which is something you can check without any technical knowledge.
Is a confidence percentage a good sign
Not on its own. A percentage attached to a single call cannot be evaluated, because any outcome is consistent with it. Percentages become meaningful only across a large scored history, which brings you back to the same question: does this tool publish its record.
My app says it uses AI. Does that change the answer
No. The test is unchanged, and it is about accountability rather than architecture. An AI system that never reports a miss is exactly as unfalsifiable as a rules based one that never reports a miss.
What if the app is right most of the time
Then it should be easy for the vendor to show you. The absence of a published record is informative in itself, because a tool with a good record has every commercial reason to display it.
Is this medical advice
No. This is wellness intelligence, meaning patterns, signals and forecasts drawn from your own data. It does not diagnose, treat, or replace medical care. If something feels wrong, talk to a clinician.
Related reading
- Why Your Health Apps Disagree About the Same Body, on what to do when several tools give different answers.
- What Is a Personal Health Baseline?, on why your own normal range beats a population range.
- Why Oura and WHOOP Give Different Scores, the same problem narrowed to two devices.
Recording what happened is largely solved. Reasoning about what happens next is barely started, and the discipline that separates the second from marketing language is scoring.
If you want a tool that states its expectations before the day happens and publishes whether they held, that is what we are building. Early access opens in cohorts through 2026. Join the waitlist at nuvard.ai.
NuVARD provides wellness intelligence. It does not diagnose, treat, or replace medical care.
