You’ve seen the headline. “New study finds X.” There’s no way to tell, from the headline alone, whether X is one of the best-supported facts in modern medicine or one small study on nineteen mice.
Last week’s post made the case for why good research should be free and findable. This week is the part that discusses what to do with a study once you’ve got it. Finding it was the easy half.
This isn’t just a satisfying mental exercise. It’s what makes the whole 80/20 approach behind this newsletter actually work. Swiss Army Mum exists on the premise that a small number of high-leverage habits do most of the work, which means “is this actually worth my limited time and energy” is a question you’re asking constantly, about a supplement, a workout protocol, a skincare ingredient, whatever’s trending this week.
You can only answer that well if you can tell the difference between a mechanism that looked promising in a petri dish and a finding that’s held up across a decade of human trials.
New here? Start here:
Swiss Army Mum is a practical guide to long-term health for busy women, built on four pillars: Body, Mind, Glow, and Flow.
Not every tool. Just the right ones.
The Courtroom
A witness says she saw the whole thing happen. That’s worth something, and it’s also exactly why an entire documentary genre exists about people wrongly convicted on eyewitness testimony alone.
Now picture a documented paper trail instead: receipts, timestamps, a pattern built from records rather than memory. Better, but still open to a dozen alternative explanations and interpretations.
Finally, picture a controlled forensic test, run under known conditions, repeatable by anyone who wants to check the work. That’s a different category of proof entirely, and nobody would treat the eyewitness and the forensic result as equally convincing.
Health claims work exactly the same way, except almost nobody applies the courtroom instinct to them. The rest of this post gives you the concepts and the ladder. Once you have both, you genuinely can't unsee it in a headline again. A study can mean an eyewitness account or it can mean a forensic result, and knowing which one you're looking at determines how much weight a health claim deserves.
One scope note before we go further. Everything in this post, the vocabulary, the ladder, the way evidence gets ranked, describes clinical and nutritional research specifically, the kind of study that ends up shaping health advice. Other fields run on different conventions entirely. None of that makes those fields less rigorous. It just means the framework here is scoped to the questions this newsletter actually asks, not a universal theory of what counts as good evidence everywhere.
Three Ways Health Evidence Gets Made
Mechanistic research happens in cells in a lab dish (technical term: in vitro, literally “in glass”) or in animals (in vivo, “within the living”). It answers how or why something might work at a biological level. Think of it as checking an engineering blueprint.
Observational research means the researcher doesn’t assign anything. She watches what people are already doing, or what’s already happened to them, and records the outcome. Two groups of women, one that happens to take a particular supplement and one that doesn’t, followed for years and compared. Its strength is scale and realism: real people, real lives. Its weakness is that the two groups were probably different before the study even started, in ways the researcher measured and plenty of ways she didn’t (a problem researchers call confounding), so a gap in outcomes might be caused by the thing being studied, or by something else entirely that happens to travel alongside it.
Interventional research means the researcher actively assigns the exposure, and ideally does it at random (randomization). Randomly assigning who gets what is what makes the two groups statistically identical going in, including in every way nobody thought to measure. That’s what lets you say a treatment caused an outcome, rather than merely showed up next to one.
A few more words travel alongside these three, and they don’t all apply evenly. A study is prospective if it follows people forward from today, and retrospective if it looks backward at records or memory that already exist. Randomization applies to interventional research. Blinding, keeping participants and researchers unaware of who got the real treatment versus a placebo, is overwhelmingly an interventional concept too.
In short: mechanistic evidence shows a pathway is plausible, observational evidence shows what already happened without controlling for other explanations, and interventional evidence, ideally randomized, is the only kind that can show what actually caused what.
The Ladder
Sit these study types on a ladder and you get the familiar hierarchy of evidence, ranked from an educated hunch up to the strongest evidence for what actually happens in real people.
Mechanistic research doesn’t sit on this ladder as a low rung. It sits underneath the whole structure, generating hypotheses worth testing rather than proving anything happens in a living person.
Near the top of that ladder, randomized controlled trials get called the gold standard so often the phrase is almost a cliché, worth knowing exactly why. It's the design every other approach gets measured against, because randomly assigning who gets what is the one thing that controls for every difference between groups, including the ones nobody thought to look for.
Above even that sits the systematic review and, at the very top, the meta-analysis: not a new study in its own right, but a careful combination of everything already published on a question. A meta-analysis pools the results of many trials into one estimate, which smooths out the kind of fluke result a single small study can produce. That combining power is also its limit: a meta-analysis is only as sound as the individual studies feeding into it.
Another variable to keep in mind in any study is who funded it. Who paid for a study shouldn’t shape what gets studied, published, and how the results get framed. However, a 2017 Cochrane review of industry-sponsored drug and device trials found them 27% more likely to report results favoring the sponsor's product and 34% more likely to reach a favorable conclusion, and the gap couldn't be explained by the studies simply being lower quality.1 It's not unique to pharma either: an analysis of 111 nutrition studies on beverages found industry-funded research was more than seven times more likely to reach a favorable conclusion than independently funded research on the same question.2
The 80/20
You don’t need to memorise the whole ladder to protect yourself from most bad science communication. Three questions do almost all of the work, even if this is the only paragraph you read.
Is this actually a claim about cause, or just two things moving together? Ice cream sales and drowning deaths both rise every summer, and nobody thinks ice cream causes drowning, they’re both just driven by hot weather. That’s correlation: two things happening alongside each other. Causation means one thing actually produces the other, and observational data on its own almost never proves that, no matter how tight the pattern looks.
Was this actually tested in people, or just in cells or animals? A compound that kills cancer cells in a petri dish, or reverses a condition in mice, is genuinely useful science. It’s just not the same claim as a finding that it does anything in a human body, and “may show promise in early lab studies” has a way of becoming “cures cancer” by the time it reaches a headline.
Who funded it, or who benefits if you believe it? A study funded by the company selling the product isn’t automatically wrong, but it’s a reason to look harder at the methods before trusting the conclusion, and it’s exactly the kind of detail a headline, or a meta-analysis, can bury without ever technically lying to you.
Why Nutrition Science Almost Never Gets to “Proven”
Most nutrition claims sit lower on this ladder than people expect, and that fact alone gets mistaken for bad science. It usually isn’t. It’s a structural ceiling, and it’s worth understanding once so you stop being surprised by it.
You cannot blind a diet the way you can blind a pill. Someone eating a Mediterranean-style diet for two years knows she’s eating a Mediterranean-style diet; there’s no version of that trial where the participant doesn’t know what’s on her plate. Long-term controlled feeding studies, where researchers actually control every calorie a person eats for years, are staggeringly expensive and logistically brutal, which is why almost none exist at the timescale that would settle most nutrition debates. Most nutrition research instead relies on people accurately remembering and reporting what they ate (a food frequency questionnaire), a method with well-documented gaps between what people report and what they actually consumed. If you’ve ever tried to recall exactly what you ate last Tuesday, you already understand the problem.
None of that makes nutrition science untrustworthy. It means the ceiling on what any single study can prove is lower than it is for, say, a drug trial, and claims that outrun that ceiling deserve extra scrutiny rather than automatic belief. A few other patterns are worth watching for regardless of topic: a mechanism presented as if it were proof in people, a single small study standing in for a whole body of evidence, a headline stating correlation as if it were causation, and a claim with no mention of who funded the work or who profits if you act on it.
Supplements sit in a slightly different spot than whole diets, though, and it’s worth knowing why before we get to today’s example. You can’t blind a Mediterranean-style eating pattern, but you absolutely can blind a capsule. Give one group real creatine and another an identical-looking placebo, and neither the participants nor the researchers running the trial need to know who got which. That’s exactly why a single compound like creatine can rack up decades of clean, well-controlled RCTs the way a drug would, in a way a whole diet never quite can.
Putting It Into Practice
Harnessing AI Tools for Literature Search
Time to put all of this to work on something real. Last week's post introduced the AI tools that make searching real scientific literature fast, Consensus chief among them. This week's ladder is what you use to judge what those tools hand back to you. So let's run both at once, on a real claim, searched with AI and graded against the framework above, so you can see exactly where a tool like Consensus earns your trust and exactly where it still needs you to think for yourself.
Before asking anything complicated, it’s worth seeing Consensus do the thing it’s actually built for: answering one clean question. I asked it something about as settled as health questions get.
That’s the tool at its best. One question, one meter, a fast, sourced answer to something with an overwhelming, settled body of evidence behind it. No filtering needed, no nuance to lose, because there isn’t much nuance to the answer in the first place.
You can read the whole report here:
A More Interesting Question
Ask Consensus something with more than one part and it handles it better than you might expect, at least at first glance.
Creatine gets marketed for two very different things right now: physiological benefits (muscle synthesis, recovery, strength) and mental ones (memory, focus, mood). Those two claims don’t sit in the same place on the ladder from earlier in this post, so I wanted to see what happens when you ask about both at once.
Credit where it’s due. Consensus didn’t flatten the compound question into mush. It split the answer into two proper tables, physiological benefits and mental and cognitive benefits, each with its own evidence-strength rating per outcome. Muscle strength in postmenopausal women came back Strong. Most of the cognitive outcomes, memory, processing speed, mood, came back Moderate.
However, the headline sitting above it, the one figure most people actually read, the meter at the top, gave a single blended answer: 75% yes, 25% possibly. That number averages a decades-deep, settled finding together with a newer, thinner one, and if you stop at the first line, you’d never know one of those claims is standing on far more evidence than the other. The nuance exists. It’s just buried under the number designed to save you from reading the detail at all.
Two more limits showed up. Throughout, not once does Consensus mention who funded any of the studies it cites. Industry-funded research and independently funded research get treated identically, with no flag either way. And when I wanted dosing information too, that meant an entirely separate question afterward, asked and answered on its own.
Before running any of these, I filtered for quality, the same way I’d want you to think about filtering, whether you’re using Consensus or reading a study yourself. Publications from the past ten years, unless there’s an older paper that’s genuinely foundational to the field, because ten-year-old data on a fast-moving topic can miss a lot, but throwing out the paper that started the whole line of research isn’t smart either. Journals ranked Q1 or Q2 by Scimago, a measure of a journal’s actual standing in its field rather than just its name recognition. Preprints excluded, since those haven’t been through peer review yet, whatever their eventual quality turns out to be.
Consensus Is A Great Tool But You Can Make it Better
Consensus’s tables are good. Its headline number still blends a strong finding and a shaky one into a single score, never checks funding, and can’t take a follow-up question in the same breath. A well-built prompt can do all three.
Build Your Own Research Brief
Once you know the ladder, you don’t have to apply it by hand every time, and you don’t have to restrict yourself to one clean question at a time either. You can hand a whole literature search to an AI tool, filters and all, and get a structured brief back instead of a single meter, one that keeps evidence-strength ratings separate instead of averaging them together, checks who funded the research, and takes a follow-up question in the same pass.
I’ve written and tested exactly that prompt. It’s the newest entry in the SAM Prompt Vault, which will be available to paid subscribers below. If you’re not one yet, this is exactly what the paid tier is for, not more science, just the ready-made instruments built on top of it.
Paste this into your AI assistant with web search enabled, filling in your actual question. Ask everything you want to know in one go, benefits and dosing and anything else, rather than working through it one question at a time.
If you’re using Claude or GPT, run this in Research mode rather than a quick chat response. Checking conflicts of interest paper by paper and hunting down decades-old foundational work that a recency filter excluded both take genuine multi-step digging, several searches, cross-checking, following citation trails, not a single quick lookup, and that’s exactly what Research mode is built to do.
The Same Creatine Question, According to The Research Brief Prompt
Rather than searching from scratch, I fed it the Consensus report and asked it to audit them directly and expand the analysis. Here’s the full brief (TL;DR stands for “Too long, didn’t read” an extreme summary):
Before going any further, everything above is still an AI-generated summary, an unusually careful one, checked against its own gaps twice over, but a summary all the same. I didn't personally read every study both reports drew on, and you shouldn't feel obligated to either. What I did do, and what I'd genuinely recommend before you act on anything here, is open the actual paper behind whichever specific finding you're going to base a real decision on, especially the ones with something at stake. Going to start taking a supplement because of a claim in this post? Spend the five minutes it takes to read the abstract of the study behind it yourself.
Same rule from last week: treat this as a first draft. AI research tools, even a well-built one, are for finding and organising the evidence faster. They're not a substitute for checking that the claims behind a specific paper are actually true. Check your sources.
One of the strongest-looking cognition papers Consensus cited without comment, a 2024 meta-analysis of 16 trials3 turned out to have a documented statistical problem. A formal published correction and a separate commentary flagged a “double-counting” error in how the studies pooled their data, and the EU’s own food safety authority raised the same concern in a 2024 opinion. None of that shows up anywhere in Consensus’s report.
The funding check turned up something more interesting than a simple pass or fail. The single strongest source behind the headline claim, that creatine reliably builds strength in postmenopausal women, turned out to be a 2026 meta-analysis4 whose lead author chairs a creatine-industry-funded advisory board, with the paper’s own publication costs partly covered by a creatine manufacturer. On its own, that’s a real reason for caution. But the brief didn’t stop there, it checked whether the finding holds up independently, and it does: several separately funded trials, backed by public health research bodies with no supplement-industry connection, reach the same conclusion. The muscle claim survives (phew!).
What doesn’t fare as well: some of the newer, more specific claims, about sleep, about mood in healthy women, about effects across the menstrual cycle, trace back almost entirely to one research network with disclosed industry ties, with no independent replication yet.
One more thing Consensus’s ten-year filter simply couldn’t have found: the actual foundational dosing protocols this entire field still runs on come from two papers published in 1992 and 1996, decades before anything in Consensus’s results. The “5 grams a day” recommendation you’ll see traces back to that original work.
Conclusion
So, as far as creatine is concerned, for what it’s worth: the strength and muscle claim is as solid as supplement science gets, genuinely well-established, independently replicated, worth the modest cost and the daily habit. The bone claim doesn’t hold up, the biggest and best-designed trial found nothing, no matter how often it gets bundled in with the strength pitch. The mental and mood claims are real but young, promising enough to watch, not yet strong enough to be the reason you start taking it.
This is what the ladder is actually for. Not memorising eleven rows of a table, but building the reflex to ask which rung a claim is standing on, and who’s holding it up.
Do that consistently, and you’ll spend your limited time and money on the handful of things that have actually earned it, which was the whole point of this newsletter long before AI made the research easier to find.
Thank you
Thank you for reading, sharing, and supporting this work. Whether you’ve been here since the beginning or just found Swiss Army Mum, I’m glad you’re here.
Building a life with more intention takes a village. If something resonated, I’d be grateful if you forwarded this to someone who might need it, or hit the ♥️ or ↻ Restack button. It helps more people find this space.
This is educational, not personal medical advice. Your biology, history, and context matter. Work with a qualified healthcare professional.
References
Lundh A, Lexchin J, Mintzes B, Schroll JB, Bero L. Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews. 2017;(2):MR000033. https://doi.org/10.1002/14651858.MR000033.pub3
Lesser LI, Ebbeling CB, Goozner M, Wypij D, Ludwig DS. Relationship between funding source and conclusion among nutrition-related scientific articles. PLoS Medicine. 2007;4(1):e5. https://doi.org/10.1371/journal.pmed.0040005
Xu C, Bi S, Zhang W and Luo L (2024) The effects of creatine supplementation on cognitive function in adults: a systematic review and meta-analysis. Front. Nutr. 11:1424972. https://doi.org/10.3389/fnut.2024.1424972
Naddafha, S., Antonio, J., Kreider, R. B., & Stout, J. R. (2026). Creatine monohydrate for lean mass, strength, and bone density in postmenopausal women: a systematic review and meta-analysis. Journal of the International Society of Sports Nutrition, 23(1). https://doi.org/10.1080/15502783.2026.2668435







