The Peptide AppEvidence review5 min read

Combinations and evidence

Each ingredient in a stack needs its own evidence grade

Formal grading finds evidence certainty varies widely between parts of one protocol, so grade each ingredient. The exact combination is rarely tested.

By , chemist and biochemist

Disclosure: Jay is a co-founder of The Peptide App. This review discusses the studies cited below; it is not a comprehensive live trial registry or treatment recommendation. Development and regulatory status can change. The app’s tools organize records and arithmetic and do not validate a research product.

Watercolor illustration of five small glass vials of different shapes in a row on a wooden tray, each under its own magnifying glass.
On this page

Key facts

QuestionDirect answer
If every ingredient in a stack has some evidence, does the stack work?No, that does not follow. When researchers formally grade multi-component protocols, certainty of evidence routinely varies enormously between components of the same protocol, from high-certainty randomized data down to no direct evidence at all [2]⁠[3].
Has the exact combination usually been tested together?Rarely. Most stack evidence is evidence for pieces tested separately in different populations; trials that test a combination directly, such as adding a second agent on top of standard therapy, are the exception [4].
Is a mechanistic story ("A raises one pathway, B lowers a competing one") the same tier as trial evidence?No. Formal grading frameworks treat biological plausibility as a distinct category, lower than outcome data from controlled trials [5]⁠[7].
Does adding more components mean more benefit?Not established. Guideline panels that build multi-part protocols assemble them from evidence graded per component, not from an assumption that stacking adds benefit [7].
Can an ordinary person grade their own stack?Yes. List each component, find its best available human evidence, assign it a tier, then work out what share of the monthly spend sits in the weakest tier.

8 sources cited. View sources

Why does stack evidence have to be graded ingredient by ingredient?

Stack evidence has to be graded ingredient by ingredient because "there is a study" and "there is a trial that changes what you should believe" are different claims. Formal frameworks sort evidence into explicit tiers instead of treating all published research as interchangeable support.

An umbrella review of ultra-processed food research shows the method. Reviewers classified pooled findings as convincing, highly suggestive, suggestive, weak, or no evidence, and out of 45 pooled analyses only a minority reached the higher tiers, even though associations turned up across dozens of health outcomes [6]. The topic is beside the point; the structure is the lesson. A large body of published findings still sorts into a small number that clear the bar for confidence.

The same logic applies to any list of things being combined. A falls-prevention review noted that many interventions have been studied, and that effectiveness, along with practical factors and patient preference, should determine which ones get used [1]. "Many interventions exist" does not mean they are equally supported, and inclusion in a protocol does not mean a component has been tested at the dose or in the combination being sold.

How much does evidence certainty vary within one treatment category?

In eczema, a network meta-analysis of systemic treatments covering 149 trials and 75 interventions found high-certainty evidence for some options and far weaker support for others [2]. The best-supported systemic options ranked among the most effective across several patient-important outcomes [2].

A companion review of topical eczema treatments, spanning 219 trials and 68 interventions, found a similarly wide range. Some agents showed high-certainty benefit across most outcomes measured, others moderate, others less [3]. These are not fringe therapies compared with snake oil. They are legitimate, studied options sitting at different points on the same certainty scale, inside the same treatment category.

Tardive dyskinesia guidance shows the same pattern from a different angle. Reviewers found that switching to a different antipsychotic carried the strongest evidence among the options examined, other proposed co-interventions had weaker support, and the overall recommendations were graded, not presented as uniformly settled [8]. A reader skimming a list of tardive dyskinesia "treatment options" would have no idea, without the grading, that one option rests on materially firmer ground than its neighbors.

What does real combination evidence look like?

Real combination evidence tests the bundle itself in a defined population and reports a number. Evidence for ingredient A, for ingredient B, and for A plus B are three separate questions, and stack marketing skips the third, which is answered far less often than the first two.

One of the clearer examples comes from inflammatory bowel disease research. A subgroup analysis found that combining a standard drug (5-ASA) with probiotics showed a higher odds ratio for inducing remission in mild-to-moderate ulcerative colitis (OR 2.35) than probiotics alone across the broader analysis (OR 2.00) [4].

That is a real test of a combination, in a defined population, with a number attached. It is also unusual. The guide to testing whether a drug combination works sets out what that kind of evidence requires.

How do clinical guidelines build multi-part protocols?

Clinical guidelines build multi-part protocols by assembling separately graded evidence for each piece, not by testing the bundle itself. A physical therapy guideline for hip and knee osteoarthritis constructs four distinct treatment profiles, each combining education with a different intensity of supervised exercise, and uses the GRADE Evidence-to-Decision framework to weigh the evidence behind each component [7]. Even that guideline is explicit about what is graded and what is judgment.

A scoliosis-exercise meta-analysis likewise found that effects varied by intervention type and duration, so "the exercise approach works" was never a single, uniform verdict [5]. Most bundled protocols follow the guideline pattern. The analysis of peptide blends sold today applies the same test to commercial blends.

How do you grade the evidence behind your own stack?

Grading your own stack means listing every component and its monthly cost, finding the single best piece of human evidence for each, assigning a tier, and summing cost by tier. The tiers, from strongest to weakest:

  • Randomized trial at a matching dose
  • Small or pilot controlled trial
  • Observational or associational only
  • Mechanistic or preclinical only
  • No direct evidence

Take a hypothetical five-part stack that costs $300 a month. If one component costs $40 and has a randomized trial behind it at a matching dose, that component earns the top tier. If the other four components, totaling $260, rest on mechanistic reasoning or animal data with no human outcome trial, then roughly 87% of that hypothetical spend sits in the lowest evidence tiers, however confident the marketing copy sounds about the one well-supported piece.

The output is a percentage, not a verdict, but it is a percentage nobody selling the stack computes for you. Five-compound longevity stacks are a real-world case worth running through the exercise.

What is still unknown about stack combinations?

For most protocols circulating without citations, direct combination-effect data, the kind shown for probiotics added to standard drug therapy [4], do not exist. Guideline-level bundles, even carefully constructed ones, are typically graded component by component, not tested as combinations [7].

That gap does not mean a stack does nothing. It means the accurate answer to "is the combination doing anything beyond its individual pieces" is usually "unknown," and that unknown is the part a buyer is often paying the most to guess about.

Sources

  1. Pillay J, Gaudet LA, Saba S (2024). Falls prevention interventions for community-dwelling older adults. Syst Rev. PMID 39593159

  2. Chu AWL, Wong MM, Rayner DG (2023). Systemic treatments for atopic dermatitis (eczema). J Allergy Clin Immunol. PMID 37678577

  3. Chu DK, Chu AWL, Rayner DG (2023). Topical treatments for atopic dermatitis (eczema). J Allergy Clin Immunol. PMID 37678572

  4. Estevinho MM, Yuan Y, Rodríguez-Lago I (2024). Efficacy and safety of probiotics in IBD. United European Gastroenterol J. PMID 39106167

  5. You MJ, Lu ZY, Xu QY (2024). Effectiveness of physiotherapeutic scoliosis-specific exercises. Arch Phys Med Rehabil. PMID 38719166

  6. Lane MM, Gamage E, Du S (2024). Ultra-processed food exposure and adverse health outcomes: umbrella review. BMJ. PMID 38418082

  7. van Doormaal MCM, Meerhoff GA, Vliet Vlieland TPM (2020). A clinical practice guideline for physical therapy in patients with hip or knee osteoarthritis. Musculoskeletal Care. PMID 32643252

  8. Ricciardi L, Pringsheim T, Barnes TRE (2019). Treatment Recommendations for Tardive Dyskinesia. Can J Psychiatry. PMID 30791698

Last updated

Junaid “Jay” Spall

Written by

Chemist and biochemist. Co-founder and author, The Peptide App.

Jay is a chemist, biochemist and entrepreneur whose work connects scientific research with consumer health products. He has held Chief Science Officer and product development leadership roles and previously served as Chief Revenue Officer at Minicircle.

Profile and articlesLinkedIn

Keep reading

The Peptide App

Track protocols, doses, and reconstitution in one place.

Save your calculations, set reminders, log doses, and keep outcome notes — free to start.

Download on the App Store