Search the published literature for BPC-157 and tendon healing and you will find something that looks, at a glance, like a mature field. Dozens of papers. Three decades of continuous output. Consistent findings across injury models — transected Achilles tendon in rats, transected medial collateral ligament, crushed muscle, damaged gastric mucosa. Reported effects in the same direction almost every time. For a compound that has never been through a published controlled human efficacy trial, that is an impressive-looking body of work.
Then you start reading the affiliations, and the field gets much smaller.
The overwhelming majority of this material comes from one research tradition: a programme based in Zagreb, together with laboratories and collaborators connected to it, publishing from the 1990s through the 2010s and beyond.[1] The models are theirs. The scoring conventions are theirs. The experimental design that produced the first striking result is, in recognisable form, the design that produced the fortieth. What looks from a distance like independent convergence is closer to a single sustained argument, delivered many times.
This is worth saying carefully, because it is easy to hear as an accusation, and it is not one. Long programmes of work by a single group are entirely normal in science, and often they are how difficult questions get answered at all. Someone has to care enough about a compound to spend thirty years on it. The problem is not that the work exists. The problem is what a reader is entitled to conclude from it, and the answer is: considerably less than is routinely concluded.
What the rat models actually show
Start with what was measured, because a great deal gets lost between the bench and the summary. In the tendon work, the typical preparation is a surgical transection of the Achilles tendon in a rat — commonly Sprague-Dawley or Wistar — followed by assessment over days to weeks. The outcomes reported include histological grading of the healing site, biomechanical measures such as the load at which the repaired tendon fails, functional scoring of gait or limb use, and descriptions of vascular patterning around the injury.
Those are reasonable endpoints. They are also, with the partial exception of load-to-failure, judgement-dependent. A histological damage grade is a human being looking down a microscope and assigning a number from a rubric. Functional scoring is a human being watching a rat walk. Both are entirely standard, and both are exactly the kind of measurement that blinding exists to protect.
Reporting on that point is inconsistent across this corpus. Some papers describe randomisation and blinded assessment; many older ones do not describe either, which was common practice at the time and is not evidence of anything improper. It does mean that for a substantial share of the literature, a reader cannot rule out the most ordinary explanation in experimental biology: that people who knew which animal received which treatment scored those animals slightly differently, without intending to.[3]
Cohort sizes compound this. Six to twelve animals per arm is the norm here, as it is across much of preclinical research. A study that size can detect a very large effect. It cannot reliably distinguish a moderate effect from noise, and it is unusually vulnerable to a couple of animals in either direction moving the result.[4]
Why single-lab dominance is a methodological problem
Replication is not a ritual. It exists because any single laboratory carries a set of shared, largely invisible commitments: how the animals are housed and handled, which supplier they come from, how the injury is made, when the endpoint is assessed, which rubric is used to score it, what counts as an outlier. Every one of those choices is defensible. Every one of them can also be quietly wrong in a way that pushes results consistently in one direction.
Repeating an experiment within the same laboratory does not test those commitments — it inherits them. Forty papers from one tradition can be forty confirmations that the tradition is internally consistent, which is a much weaker claim than forty confirmations that the effect is real. The errors, if there are any, are correlated. Averaging correlated errors does not cancel them out; it just makes them look more precise.
The question is not whether these results are real. It is whether this literature, as constructed, could have told us if they weren’t.
Preclinical research has learned this expensively. In the early 2010s, research groups at two pharmaceutical companies published widely discussed accounts of trying to reproduce landmark preclinical findings in-house and failing to confirm the majority of what they attempted.[2] A large, multi-year replication project in cancer biology, reporting in the early 2020s, found that replicated effects were substantially smaller on average than the originals, with a meaningful fraction not reproducing at all.[5] None of that work concerned peptides. All of it concerned exactly this structure: influential findings, small studies, incomplete reporting of randomisation and blinding, and very little independent repetition before the result entered general circulation as established.
The peptide literature has not been through that reckoning. It is, in several respects, more exposed to it than the fields that have.
What independent replication would actually require
It is easy to call for replication and much harder to specify one that would settle anything. A serious attempt on the tendon question would need several features that are largely absent from the existing corpus.
It would need to be run by investigators with no training lineage or collaborative connection to the original programme. It would need a protocol registered before the animals were ordered, specifying the primary endpoint in advance — one endpoint, chosen and committed to, not a panel of measures from which the most cooperative can be selected afterwards. It would need randomised allocation and outcome assessors who genuinely do not know which group they are scoring. It would need a sample size justified by a power calculation against a pre-specified effect size, which for anything short of an enormous effect means considerably more animals than this field typically uses. And it would need to be published whichever way it came out, which is the requirement that quietly does the most work, because null results in this area are difficult to place and rarely rewarded.[3]
Designs of exactly this kind exist. Stroke research, after a long run of compounds that protected neurons in rodents and did nothing in people, developed the multicentre preclinical trial: randomised, blinded, protocol-registered animal studies run across several laboratories simultaneously, powered like clinical trials and reported like them.[6] They are slower and more expensive than a single-site study, and they have repeatedly produced more modest answers than the literature that preceded them. That is the point of them.
Where this leaves a reader
Nothing above shows that BPC-157 does not do what the rat papers report. That is a claim the available evidence cannot support either, and it would be its own kind of overreach to make it. Single-source literatures are sometimes vindicated. Occasionally one group is simply the only one paying attention to something real.
What can be said is narrower and duller. There is a substantial body of rodent work, concentrated in one research tradition, using small cohorts, with variable reporting of the safeguards that protect against the most common sources of error, and very little independent repetition. There is no published controlled human efficacy trial. Between those two sentences sits a large amount of confident writing that is not supported by either.
The honest position is that this is an unresolved question that has been treated as a resolved one, and that the thing standing between it and resolution is not more papers. It is different papers, by different people, designed to be capable of coming out the other way.
References and further reading
- The primary rat tendon and ligament literature on this compound, published across the 1990s, 2000s and 2010s, largely from a research programme based in Zagreb and from laboratories affiliated with it. Described here at the level of models, endpoints and reporting practice rather than by individual paper.
- Two commentaries from pharmaceutical industry research groups, published in 2011 and 2012, reporting that in-house teams were unable to confirm the majority of landmark preclinical findings they attempted to reproduce.
- The ARRIVE guidelines for reporting animal research, first published in 2010 and revised in 2020, which set out minimum reporting expectations including randomisation, blinding, and justification of sample size.
- The general statistical literature on power in small-sample animal research, including the relationship between low power, inflated effect-size estimates, and reduced probability that a reported positive finding is true.
- A large multi-year replication project in cancer biology, reporting in the early 2020s, which attempted a set of high-profile preclinical experiments and found replication effect sizes substantially smaller than the originals.
- The multicentre preclinical trial literature developed principally in stroke research: randomised, blinded, protocol-registered animal studies run across multiple laboratories and reported to clinical-trial standards.
We describe sources at this level deliberately. We do not print identifiers for papers we have not read in full, and we do not manufacture the appearance of precision. Readers wanting the primary material can find it through the models and dates given above.