"Science says …" – who is that actually, and who paid for it?

Wissenschaftlerin macht sich im Labor Notizen

© Adobe Stock / peopleimages.com

This article was translated using AI.

How studies are actually created, what they can achieve, and where the limits of research and practical experience lie

Highlights

  • In 2023, the industry contributed more than two-thirds of all expenditure on research and development in Germany. At universities, around every seventh euro of third-party funding came directly from the commercial sector (Destatis 2025a, Destatis 2025b).
  • Industry-funded studies systematically more often reach results that benefit the sponsor – in the Cochrane review of 4,583 studies as well as in veterinary medicine (Lundh et al. 2017, Wareham et al. 2017a).
  • In veterinary therapy studies with pharmaceutical funding, positive results were reported in 56.9 percent of cases; without pharmaceutical funding, only in 29.1 percent (Wareham et al. 2017a).
  • Only 14.3 percent of these studies contained any sample size calculation, and only 40.5 percent had a pre-specified primary endpoint – with a median of five measured outcomes (Wareham et al. 2017b).
  • Reviewers in the peer-review process found an average of only 2.58 out of nine intentionally included major methodological errors in a controlled experiment (Schroter et al. 2008).
  • In 2023, for the first time, more than 10,000 specialist articles were retracted worldwide (Van Noorden 2023).
  • On average, it takes around 17 years from a research discovery to its application in practice (Morris et al. 2011).

 

The sentence that ends every conversation

“That is not scientifically proven.” Anyone who hears this sentence has already lost. Not because it might be wrong, but because it is constructed in a way that makes it impossible to answer. It ends discussions instead of opening them. And on the web, it is now used as matter-of-factly as a moralizing finger once was.

On the other side stands its mirror image: “Science says.” This is usually followed by a surprisingly specific claim, supposedly drawn from some study whose title no one mentions and whose methodology no one has read. Both sentences have the same problem. They treat science like a person with an opinion – an authority that says something and is therefore always right.

But science is not a person. It is a process. A very good, very laborious, but in some parts also broken process, conducted by people who have to pay rent, have careers, and are paid by someone. Anyone who wants to understand what a study's statement is worth must not only read the headline and conclusions but understand how it came to be.

And to be clear from the start: This is not a text against science. Almost everything written here in terms of criticism comes from science itself – from The Lancet, the British Medical Journal, Cochrane reviews, and veterinary journals. The industry knows its problems very well. It just hasn't solved them yet.

What the double-blind study can really do

Before the criticism comes, one must understand why controlled studies were invented in the first place. They are not an academic gimmick, but the answer to a problem that affects every human being – not just in the stable or the veterinary clinic: our perception is systematically unreliable when judging effectiveness.

This has several reasons that all act simultaneously. First, most complaints get better on their own. Anyone who treats a horse exactly when it is at its worst will almost always see an improvement – even without treatment. Statisticians call this “regression to the mean.” Second, we see what we expect. Someone who has spent 80 euros on a supplement judges the coat differently than someone who hasn't fed anything extra. Third, with a new measure, almost never just one thing changes: someone who introduces a supplement often simultaneously cuts the concentrate feed, pays more attention to the roughage, and moves the horse more consciously and regularly. In the end, no one knows which of the five changes worked.

The randomized, controlled, double-blind study was introduced precisely to counteract this. Randomization – the random assignment of animals to groups – ensures that the groups are similar on average and that healthier horses don't end up in the treatment group while sicker ones end up in the control group. The control group shows what would have happened on its own without treatment. Blinding prevents the person treating and the person evaluating from unconsciously seeing what they want to see – a factor referred to as the placebo or nocebo effect. And the pre-specification of what will be measured prevents someone from picking the one value out of twenty that looks good after the fact and building a result from it.

This is a magnificent invention, and it has saved countless lives – both human and animal. No serious practitioner should downplay that. The only question is: how many of the studies used for arguments online are actually built this way? And who paid for them?


Terms frequently encountered regarding “studies”

Control group — A group of animals that does not receive the investigated treatment but is otherwise kept and assessed in exactly the same way. It shows what would have happened even without treatment. Without a correct control group, you measure the passage of time and book it as an effect, which may have nothing to do with the method at all.

Placebo — A sham preparation without an active ingredient that looks, smells, and tastes just like the real thing. In horses, a placebo doesn't act on the animal, but it does act on the human who feeds it and then judges it. That is exactly why it is needed: whoever administers something expects an effect.

Blinding / Double-blind study — Blinded means: the evaluator does not know which animal received which treatment and which received a placebo. Double-blind means: even the person treating or feeding does not know. This eliminates the strongest source of error: human expectation.

Randomization — The assignment of animals to groups by chance instead of by gut feeling. It prevents healthier, younger, or calmer horses from unconsciously ending up in the treatment group. “We selected the groups carefully” is the opposite of randomized.

Cross-over design / Latin square — Each animal receives all treatments one after another in a different order. This saves animals because every horse passes through every group over the course of the study. This way, one can form four groups with four horses and end up with 16 measurements. The catch: there must be a sufficiently long washout phase between cycles, otherwise the previous treatment may have a carry-over effect; therefore, distortion of results is not always excluded in this study design. And the number of animals remains small, even if many measured values appear in the table at the end; results should thus be interpreted with caution.

Primary endpoint (Main outcome measure) — The one variable determined before the start of the study to define whether the treatment works. Everything else is a secondary finding. Anyone who does not specify a primary endpoint can later pick the most attractive one from many measurements and declare it the result.

In vitro and in vivo — In vitro means “in glass,” i.e., in a test tube, on cell cultures, or on tissue samples. In vivo means “in the living animal.” Many efficacy promises for supplements are based on in-vitro findings. However, the fact that a substance influences a cell in the lab says nothing about whether it actually gets from the intestine into the blood and then to its target in the horse.

 

How a study is created – and who pays for it

The most common misconception about research is that it is paid for neutrally, for example, by tax money. One imagines a professor investigating a question of interest using public funds and reporting the results openly at the end. This image is about fifty years old.

In fact, around 129.7 billion euros were spent on research and development in Germany in 2023. Of this, the industry contributed more than two-thirds with 88.7 billion euros (Destatis 2025a). At the universities themselves, it looks better at first glance – there, of 10.7 billion euros in third-party funding, 3.3 billion came from the federal government and 3.2 billion from the German Research Foundation (DFG). But 1.54 billion euros, or roughly one in seven third-party euros, came directly from the commercial industry (Destatis 2025b). And that is only the part officially recorded as industry money; benefits in kind, products provided free of charge, sponsored doctoral positions, and congress funding do not appear in these statistics.

Why is this? Because research is expensive – considerably more expensive than most people suspect. The German Research Foundation estimates 6,800 euros for a doctoral position and 7,350 euros for a postdoc position per month (DFG 2026). These are not salaries, but the full costs that a project must budget for a single person – around 82,000 or 88,000 euros a year. A three-year feeding study with a single doctoral position thus costs around a quarter of a million euros in personnel alone.

Then there is everything else. The experimental horses must be kept, fed, provided with veterinary care, and stabled individually for months so that it's known who ate what. In digestibility studies, every dropping is fully collected and weighed. Feed must be analyzed, blood as well, often at several points in time per animal and day. Laboratory analysis is a unit-cost business: every additional sample costs real money. An animal testing application must be written, reviewed, and approved. In the end, many journals demand a publication fee in the four-figure range so that the article is freely readable.

Who has an economic interest in paying for this? In horse nutrition, the answer is usually simple: the feed manufacturers. For a manufacturer, a study is not an academic quest for knowledge, but an investment with an expected return. No one finances an investigation whose most likely result is that their own product is useless or perhaps even harmful.

One does not need to assume fraud here. The effect is more subtle and therefore more effective. It lies in the research question: you don't compare your own product with the best alternative, but with nothing at all. It lies in the selection of outcome measures: you measure twenty parameters and report those that show something. It lies in the choice of the control group, the dosage, the duration – a study over six weeks will fundamentally never find long-term damage. And it lies in which results actually find their way into a journal and which remain in a drawer.

That this effect is real and measurable is one of the best-documented facts of research methodology. The relevant Cochrane review evaluated 75 investigations with a total of 4,583 studies. Industry-funded drug and medical device studies more frequently reached efficacy results favorable to the sponsor and more frequently reached favorable overall conclusions than independently funded ones. Particularly revealing is a side finding: in industry-funded studies, the results section and conclusion matched less often than in independent ones (Lundh et al. 2017). Translated, this means the numbers say one thing and the summary in the abstract says another – and usually, only the summary is read, not the methodology and results.

The same investigation exists for the nutrition sector. Lesser and colleagues analyzed 206 publications on soft drinks, juice, and milk. Articles funded exclusively by the respective beverage industry were four to eight times more likely to reach conclusions in favor of the payer. Of the 16 purely industry-funded intervention studies, not a single one came to a result unfavorable to the sponsor, while in the non-funded or mixed-funded studies, 7 out of 19 found an unfavorable effect (Lesser et al. 2007).

And because the objection will come that this is human medicine and food and everything is different with animals: it is not different. Wareham and colleagues examined 126 randomized controlled studies from a single year involving cats, dogs, horses, cattle, and sheep. In the group with participation or funding from the pharmaceutical industry, 56.9 percent positive results were reported. Without pharmaceutical involvement, it was 34.9 percent, and for studies with no information on funding, 29.1 percent. Additionally, the authors found that a high proportion of all studies examined had an unclear risk of bias because they were simply too poorly published to be assessed (Wareham et al. 2017a).

How such things look historically when internal documents surface decades later is shown by probably the most famous case in nutrition research. The American sugar industry financed a review paper in one of the leading medical journals in 1965. It set the objective of the work, provided literature for inclusion, and received drafts for review. The result identified fat and cholesterol as the cause of coronary heart disease and downplayed the role of sugar. Documents show that the industry association knew as early as 1954 that a low-fat dietary recommendation would increase per capita consumption of sugar by more than a third (Kearns et al. 2016). The funding was not disclosed at the time; it was not standard practice then. The “less fat” recommendation subsequently shaped two generations of nutritional advice.

This is the point where folk wisdom is more precise than any methodological critique: “Whose bread I eat, his song I sing.” Not because researchers are for sale. But because a person whose position, and thus their own living costs, depends on the sponsor providing funding again next year, makes a thousand small decisions in the study design – and while not a single one is made in bad faith, on average, they do fall in a certain direction.

Four horses, four weeks, sixteen measurements

Now to the second major topic, which is almost more important for horse nutrition than the money: the size of the studies.

An example that can show this cleanly without doing anyone an injustice: Brøkner and colleagues investigated how different carbohydrate fractions are digested in the horse and what this does to the pH value in the cecum. Four horses with a permanently implanted access to the cecum (fistulated research horses), four rations, a so-called cross-over design, seventeen days of adaptation per ration, eight days of sampling, and pH measurement every minute over eight hours were used (Brøkner et al. 2012).

This is a methodologically complex and scientifically perfectly legitimate piece of work. It is built exactly how things are built in horse nutrition, and the authors did nothing wrong. It is therefore a good example – not because it is particularly bad, but because it is typical. Four horses are a completely normal number in this line of research. Fistulated horses are extremely expensive, their number is ethically and practically limited, and no one has forty of them in the stable.

What follows from this? First of all, something positive: for the question of what happens mechanistically in the digestive tract of a horse, such studies are valuable and often the only way to learn anything at all. They show that something can happen and through which pathway.

What does not follow from this is a feeding recommendation for all horses. Four animals are four animals. They are usually Warmbloods or Thoroughbreds of middle age from a university herd, healthy, and accustomed to the facility. They are not a Shetland pony with EMS, not a 25-year-old retiree with dental problems, and not an Icelandic horse with sweet itch. With four animals, a single horse that steps out of line for some reason decides the entire result.

Added to this is a second, more subtle problem, which is the actual statistical core. If you measure many variables on a few animals, you are almost guaranteed to find a difference somewhere – even if, in reality, nothing is happening. This is not a conspiracy, but simple mathematics. A single test is usually called “significant” if the probability of getting such a result purely by chance is less than five percent. If you conduct twenty such tests, you can expect roughly one false hit purely by calculation. And this one hit becomes the headline.

That is exactly why good methodology requires two things: that you determine beforehand which variable is the decisive one – the so-called primary endpoint – and that you calculate beforehand how many animals you need to be able to find a relevant effect at all. Both are the exception in veterinary medicine. In the same sample of 126 studies that stood out regarding funding, only 14.3 percent contained a sample size calculation, and only 40.5 percent even named a primary endpoint – with a median of five measured variables per study (Wareham et al. 2017b). Systematic reviews of veterinary literature find proportions of zero to 24 percent for sample size determination (Weese 2026).

However, one must name this problem precisely, or one becomes imprecise oneself. Small studies are not worthless. Many small, cleanly conducted, and consistent studies can together support a robust statement; the value then lies in the data, not in the individual analysis, which by itself is almost always too weak to show a clinically relevant difference (Weese 2026). The problem arises where an investigation on four animals over four weeks becomes a general feeding recommendation – or an advertising slogan.


How study results are read

Median — The middle value of a series sorted by size: one half lies above it, the other below. Unlike the average, the median is not distorted by individual outliers. If an evaluation states that the median is five measured outcomes, it means half of all studies measured even more than five things simultaneously.

Significant — The most frequently misunderstood term of all. Statistically significant does not mean “important,” “strong,” or “proven.” It only means that a result of this size would rarely occur by chance under the assumption that there is truly no difference — usually in fewer than five out of a hundred cases. A significant effect can still be tiny and practically completely meaningless.

p-value — The number behind this assessment, usually given as “p < 0.05.” It does not say how likely it is that the result is correct, nor does it say how large the effect is. If you measure twenty things, you will be gifted a p-value below 0.05 purely by chance on average.

Sample size calculation and power — Before the study, it is calculated how many animals are needed to be able to find an effect of the size of interest at all. If this calculation is missing, a negative result is not meaningful: a study that is too small will not find even real effects. “No difference was found” and “there is no difference” are two different statements. If you have too few test animals, you measure values, but they have no scientific weight.

Relative and absolute risk — The favorite numerical trap of advertising. If a risk drops from two to one percent, it can be advertised as a “50 percent reduction” (relative) or as “one percentage point less” (absolute). Both are mathematically correct, but the first phrasing sounds infinitely more impressive. If only the relative figure is mentioned, the crucial information is missing.

Correlation and causality — Two things occur together; that is a correlation. That one causes the other is not proven by this. Horses that receive a certain supplement might be healthier because their owners generally feed more attentively, give more hay, and look after their animals better — and not because of the supplement itself.

Regression to the mean — Extreme states normalize over time on their own. Anyone who treats exactly when the horse is at its worst will almost always see an improvement — even with a completely ineffective remedy. This effect is the silent engine behind countless testimonials.

 

Who checks the checkers?

What remains is the authority everyone online refers to: Peer Review. The article is “peer-reviewed,” i.e., checked by experts, so it is correct. That is the short version. It's worth taking a look at what actually happens there.

Peer review means that an editorial office sends a submitted manuscript to two or three experts who read it and provide a recommendation. These experts receive no money for this. They do it in addition to their actual work, usually in the evening, usually under time pressure. A working group has extrapolated how large this unpaid effort is overall: globally, over 100 million working hours flowed into reviews in 2020, equivalent to more than 15,000 person-years. The value of reviewers' time in the USA alone was over 1.5 billion dollars. The authors emphasize that this is a clear underestimate (Aczel et al. 2021).

What does this check actually achieve? This can be measured, and it has been. In a randomized study, 607 reviewers from a major medical journal received manuscripts into which nine major and five minor methodological errors had been intentionally built. in the baseline round, reviewers found an average of 2.58 of the nine major errors. Targeted training improved this only slightly (Schroter et al. 2008). Earlier investigations summarized there came to similar results: two-thirds of the major errors built in went undetected, and a small portion of reviewers even recommended the faulty papers for publication.

This applies even more to fraud. The long-standing editor-in-chief of the British Medical Journal, who is also a co-founder of the Committee on Publication Ethics, put it very clearly in 2021: peer review is structurally unable to detect fraud because reviewers fundamentally assume the honesty of the authors. They only see the text, not the raw data and not the stable (Smith 2021).

And the selection of reviewers itself is vulnerable. Many journals ask authors to suggest possible reviewers – which makes sense for very specialized topics because the editorial office does not oversee the small specialist world. This door has been spectacularly exploited several times. In 2017, a major publisher retracted 107 papers from a single journal because the authors had provided real researcher names with self-controlled email addresses as reviewers, effectively reviewing their own work (Retraction Watch 2017).

Anyone publishing in a niche topic has good chances of being reviewed by people they know – or by those who only get to know the topic while reading.


How the scientific enterprise works

Peer Review — The evaluation of a submitted manuscript by two to three experts who are not paid for it. They check the text, not the raw data, the analysis, and certainly not the animals. “Peer-reviewed” therefore means: someone in the field has read it and had no obvious objection. It does not mean: checked, recalculated, the study reproduced and confirmed.

Third-party funding — Money that a university raises in addition to its regular budget, for example, from public research funding, foundations, or companies. In practice, there is hardly any research without third-party funding — and whoever raises it has an interest in continuing this cooperation with the payer.

Conflict of interest — Not the same as corruption. A conflict of interest exists as soon as someone benefits from a result going in a certain direction. It mostly acts unconsciously and yet measurably. Therefore, it must be disclosed, which is not always the case.

Publication bias — Studies with positive results are published more frequently than those that found nothing. Published literature is therefore systematically more optimistic than reality. Anyone who only counts how many studies found an effect is counting a pre-sorted selection.

Meta-analysis and systematic review — A systematic review searches for all studies on a question according to fixed rules and evaluates their quality. A meta-analysis additionally calculates their results into a grand total. Both are significantly more meaningful than an individual study — but only as good as the studies that go into them.

Retraction — An already published article is subsequently withdrawn due to errors, fraud, or manipulated review. Retracted papers do not disappear from the world: they are often cited for years as if nothing happened.

Consensus statement and guideline — A recommendation derived by a specialist society from the existing state of studies. It is not a new investigation, but an evaluation — and the point in time at which a finding finds its way there is usually many years after the first publication.

 

Publish or Perish

This brings us to the engine of the whole thing. Scientific careers are not awarded based on how right someone is, but based on how much and where someone publishes. Anyone without publications gets no third-party funding, no extension of their projects or employment, no professorship. “Publish or perish” is not a witty remark, but a fairly exact description of the situation researchers find themselves in.

This incentive structure rewards novelty, not truth. A positive, surprising result can be published. “We looked and found nothing” is just as valuable for knowledge but much harder to place. And a replication of a colleague's study, i.e., exactly what would actually safeguard the process, is considered unoriginal and brings little recognition - especially if the reproduction of the same study fails because the original study might have been fabricated.

The consequences of this have now been very well measured. A survey of 1,576 researchers found that over 70 percent had already tried in vain to reproduce a colleague's experiment, and more than half had even failed at their own. Around nine out of ten respondents saw a reproducibility crisis (Baker 2016). The methodological classic on this topic comes from a Stanford epidemiologist and bears the intentionally provocative title that most published research findings are false; it remains one of the most cited works of biomedical methodological critique (Ioannidis 2005).

Even open misconduct is not as rare as one would hope. A meta-analysis of surveys found that around two percent of scientists admitted to having fabricated, falsified, or altered data at least once. Up to 34 percent admitted other questionable practices – including omitting data that contradicts their own earlier results or removing individual measurements based on a gut feeling that they were probably erroneous. If one asked about the behavior of colleagues, the values rose to 14 or up to 72 percent. Researchers from medicine and pharmacology reported most frequently (Fanelli 2009).

How this manifests in the system is shown by a look at the retractions. In 2023, more than 10,000 specialist articles were retracted worldwide for the first time, over 8,000 of them by a single publisher alone, because the review process there had been systematically undermined. The retraction rate has more than tripled in a decade and is growing faster than the number of publications themselves (Van Noorden 2023).

Behind this stands a literal industry, so-called “paper mills”: companies that sell finished studies including authorship slots to order. A study published in 2026 in the British Medical Journal searched 2.6 million specialist articles from cancer research from 1999 to 2024 using a specially trained language model. Around 250,000 papers, just under ten percent, showed text patterns typical of already retracted paper mill publications (Barnett et al. 2026). Important for correct classification: these are suspected cases for verification, not proof that these works are falsified. But the magnitude of what needs to be verified at all is the real finding.

And because this is the point where the accusation of conspiracy theory arises: none of this comes from obscure sources. The editor-in-chief of The Lancet wrote as early as 2015 in his own journal, after a symposium of leading research institutions on the reproducibility of study results, that possibly half of the scientific literature could simply be untrue – and cited small sample sizes, tiny effects, methodologically invalid exploratory analyses, and open conflicts of interest as causes (Horton 2015). This is not criticism from the outside. This is the self-disclosure of the authority everyone refers to.

Why research doesn't know reality

There is yet another level that is harder to capture in numbers, but which everyone who has experienced both knows: university and the stable aisle.

Anyone researching horse nutrition at a university almost inevitably works in fragments. That is in the nature of things. A simplified scientific question is a narrow question: How does the pH value in the cecum change after a defined grain meal? How high is the insulin response to a certain feed? Anyone who asks more broadly – “is this horse better after two years?” – gets a question that cannot be answered cleanly with a reasonable research effort. So the questions are kept narrow. This is methodologically correct and yet leads to the fact that systematically no one investigates the long-term consequences of a recommendation on the living horse in a real stable.

Added to this is the social cohesion of the operation. You read the works of colleagues, you listen to them at congresses, you review each other, you cite each other. This ensures that the scientific enterprise forms its own “bubble” that often has little to do with life out in the horse stable. The practitioner who has accompanied the same two hundred horses for twenty years and therefore knows what became of a recommendation after five years does not sit at the table in this circle. His knowledge has no format in which it could reach the circles of scientists.

Knowledge transfer also often only takes place with a delay in reverse. How long it takes for scientific findings to even arrive in application has been investigated. Several independent works with different methods come to a median time lag of about 17 years between research results and routine application; depending on the topic, the span extends significantly beyond that (Balas and Boren 2000, Morris et al. 2011).

For horse nutrition, this can be shown with an example that many horse owners have experienced themselves. The connection between disturbed insulin regulation and laminitis was described in experiments on ponies as early as the 1980s. The term Equine Metabolic Syndrome was only introduced in 2002 (Johnson 2002). A first American consensus paper followed in 2010; the European consensus paper, which places insulin dysregulation at the center as a key and consistent feature, only in 2019 (Durham et al. 2019). There are therefore around three and a half decades between the first experimental hint and the European guideline. And for the majority of this period, processed grain was still recommended as feed in practice, much to the suffering of many horses.

Practice operates precisely in this gap. Anyone who advised owners of laminitic horses in the early 2000s to consistently reduce sugar and starch in the ration did not wait for a guideline but reacted to what was visible in the horse. That this recommendation is textbook knowledge today does not change the fact that it was considered highly unscientific back then. This is the honest answer to the question of why practical experience counts at all: it is often the only information available at a time when a decision must be made.

What experience can do – and what it can't

And now the part that one could omit if one only wanted to be right. Experience is not automatically better than scientific investigations. It is just prone to different errors.

The most impressive proof of this comes from human medicine. A working group compared what combined evaluations of randomized trials for the treatment of heart attacks already showed at a given point in time with what experts recommended in textbooks and reviews at the same time. The result was devastating: experts continued to recommend therapies for years whose uselessness or harmfulness had long been proven, and failed to mention effective treatments for which evidence had existed for years (Antman et al. 1992). Experience and authority were not ahead here, but lagged behind – with fatalities as a consequence.

And from the horse stable, there is a matching counterpart. For decades, it was considered good professional practice to deworm every six to eight weeks and rotate active ingredient classes while doing so. This was lived experience, done in every stable, passed from one generation to the next. It has contributed significantly to the fact that small strongyles are today resistant to several active ingredient classes, and the situation for roundworms on breeding farms looks no better (von Samson-Himmelstjerna 2012). This was corrected not by experience but by systematic investigation: today, selective deworming based on fecal samples is considered the standard, with efficacy control via the egg count reduction test (AAEP 2024). And yet, one still has to discuss this with veterinarians and stable managers in daily practice because selective deworming is accused of lacking scientific reliability.

The topic of supplements also shows what independent research is good for. In an investigation published in 2026, 40 older geldings with chronic lameness received either a commercially available oral joint preparation or a placebo over six weeks, assigned according to lameness grade, weight, and body condition. There was no effect on stride length, lameness grade, or objectively measured gait symmetry between the two groups. Quite a bit changed over time – but equally in both groups (Harbowy et al. 2026). It is exactly this time effect that one would book as an effect of the preparation in the stable without a control group. One would have been convinced that it helps. One would have recommended it further. And one would have been wrong.

The honest balance therefore looks like this: experience recognizes patterns early, works on the whole animal over years under real conditions – but cannot cleanly separate cause from chance. It has no built-in correction against its own confirmation bias, and it also doesn't see the horses whose owner looked for a different therapist because it didn't work. Science has this correction but works in fragments, short-term, on a few and mostly “standardized” animals, is slow, and allows itself to be financed, which leads to conflicts of interest. Anyone who declares one of the two as the sole standard makes the same mistake – just in different directions.

How to assess a study as a layperson

You don't have to have studied statistics to roughly categorize the value of a study statement. Seven questions usually get you surprisingly far, and the answers are almost always in the freely accessible part of the work.

  • First: Who paid? At the end of every reputable publication is a section on funding and conflicts of interest. If a manufacturer whose product was investigated is listed there, the work is not automatically wrong – but it belongs in a different trust class. If nothing is listed there, that's not a good sign: in the veterinary sample by Wareham and colleagues, works without any funding information were their own poorly assessable group (Wareham et al. 2017a).

  • Second: How many animals, and which ones? Four fistulated university horses are different from forty horses and ponies of various breeds in an open stable. The smaller and more homogeneous the group, the less the result can be transferred to your own horse.

  • Third: What control groups were there, were they real placebo groups, and was the study truly blinded? Without correct control groups, you measure time, not effect. Without blinding, you partially measure the expectation of the evaluator.

  • Fourth: How long did the study run? Six weeks say nothing about a year. Especially with feed, the relevant problems usually do not arise immediately after the adaptation phase, but after months or years of administration.

  • Fifth: What was the pre-specified primary endpoint? If a work measures fifteen variables and celebrates exactly one of them in the conclusion, it's worth looking at the results section – were the other fourteen perhaps neutral or even negative?

  • Sixth: Does the conclusion say the same as the results section? This sounds trivial, but it isn't. This gap between numbers and conclusion was demonstrably larger in industry-funded studies than in independent ones (Lundh et al. 2017).

  • Seventh: Has anyone independently repeated it? A single study is a data point, not knowledge. Only when another working group with different money comes to the same result does it become something on which a concept should be built.

  • And an eighth question, which is not a question of methodology but of integrity: Is the study named at all? Anyone who argues online with “science says” but cannot provide author, journal, or year upon request is not arguing scientifically. They are using “science” to silence their opponent – and that is exactly the opposite of what it was invented for.

How does science help us now?

Science is the best tool we have to find out if something works or if we are just imagining it. This is not a cliché, but the reason why this article begins with a chapter on the strengths of the controlled study.

But science is also an industry with sponsors, careers, time pressure, and quality control that overlooks seven out of nine built-in major errors in experiments. It only answers questions for which funding can be found, on as many animals as the budget allows, in the timeframe a doctoral project permits. And it takes an average of seventeen years for the result to arrive in the stable.

Those who know this become suspicious of both extremes. Of the sentence “that is not scientifically proven,” which acts as if everything unproven is thereby refuted – whereas “not investigated” first and foremost just means that no one was found to pay for the investigation. Opposite this stands the sentence “studies have shown,” which declares a single piece of work on four horses to be a constant of nature.

In the end, an uncomfortable but sustainable attitude remains: take both seriously, distrust both, and in each individual case, ask where the knowledge actually comes from. Anyone who claims that only “scientifically proven facts” are valid has never seriously engaged with the scientific enterprise. And anyone who fundamentally considers studies to be bought or falsified hasn't either. The horse standing before us is not interested in either position. It is interested in whether we look closely and help it to the best of our ability with all the knowledge available to us.




Sources

AAEP (2024): Internal Parasite Control Guidelines. American Association of Equine Practitioners.

Aczel, B., Szaszi, B., Holcombe, A. O. (2021): A billion-dollar donation: estimating the cost of researchers’ time spent on peer review. Research Integrity and Peer Review 6:14. DOI: 10.1186/s41073-021-00118-2

Antman, E. M., Lau, J., Kupelnick, B., Mosteller, F., Chalmers, T. C. (1992): A comparison of results of meta-analyses of randomized control trials and recommendations of clinical experts. Treatments for myocardial infarction. JAMA 268(2):240–248. DOI: 10.1001/jama.1992.03490020088036

Baker, M. (2016): 1,500 scientists lift the lid on reproducibility. Nature 533:452–454. DOI: 10.1038/533452a

Balas, E. A., Boren, S. A. (2000): Managing clinical knowledge for health care improvement. Yearbook of Medical Informatics, pp. 65–70.

Barnett, A. et al. (2026): Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study. BMJ 392:e087581. DOI: 10.1136/bmj-2025-087581

Brøkner, C., Austbø, D., Næsset, J. A., Bach Knudsen, K. E., Tauson, A. H. (2012): Equine pre-caecal and total tract digestibility of individual carbohydrate fractions and their effect on caecal pH response. Archives of Animal Nutrition 66(6):490–506. DOI: 10.1080/1745039X.2012.740311

Destatis (2025a): 7% more expenditure on research and development in 2023. Press release No. 084 of March 7, 2025. Federal Statistical Office, Wiesbaden.

Destatis (2025b): University expenditure 2023 increased by 6%. Press release No. 119 of March 27, 2025. Federal Statistical Office, Wiesbaden.

DFG (2026): DFG personnel rates for the year 2026. DFG form 60.12 – 01/26. German Research Foundation, Bonn.

Durham, A. E., Frank, N., McGowan, C. M., Menzies-Gow, N. J., Roelfsema, E., Vervuert, I., Feige, K., Fey, K. (2019): ECEIM consensus statement on equine metabolic syndrome. Journal of Veterinary Internal Medicine 33(2):335–349. DOI: 10.1111/jvim.15423

Fanelli, D. (2009): How many scientists fabricate and falsify research? A systematic review and meta-analysis of survey data. PLoS ONE 4(5):e5738. DOI: 10.1371/journal.pone.0005738

Harbowy, R. M., Robison, C. I., Tillman, I., Manfredi, J. M., Nielsen, B. D. (2026): Efficacy of an oral chondroprotective joint supplement on stride length and gait symmetry in aged geldings with chronic lameness. Animals 16(8):1230. DOI: 10.3390/ani16081230

Horton, R. (2015): Offline: What is medicine’s 5 sigma? The Lancet 385:1380. DOI: 10.1016/S0140-6736(15)60696-1

Ioannidis, J. P. A. (2005): Why most published research findings are false. PLoS Medicine 2(8):e124. DOI: 10.1371/journal.pmed.0020124

Johnson, P. J. (2002): The equine metabolic syndrome – peripheral Cushing’s syndrome. Veterinary Clinics of North America: Equine Practice 18(2):271–293. DOI: 10.1016/s0749-0739(02)00006-8

Kearns, C. E., Schmidt, L. A., Glantz, S. A. (2016): Sugar industry and coronary heart disease research: a historical analysis of internal industry documents. JAMA Internal Medicine 176(11):1680–1685. DOI: 10.1001/jamainternmed.2016.5394

Lesser, L. I., Ebbeling, C. B., Goozner, M., Wypij, D., Ludwig, D. S. (2007): Relationship between funding source and conclusion among nutrition-related scientific articles. PLoS Medicine 4(1):e5. DOI: 10.1371/journal.pmed.0040005

Lundh, A., Lexchin, J., Mintzes, B., Schroll, J. B., Bero, L. (2017): Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews 2:MR000033. DOI: 10.1002/14651858.MR000033.pub3

Morris, Z. S., Wooding, S., Grant, J. (2011): The answer is 17 years, what is the question: understanding time lags in translational research. Journal of the Royal Society of Medicine 104(12):510–520. DOI: 10.1258/jrsm.2011.110180

Retraction Watch (2017): A new record: Major publisher retracting more than 100 studies from cancer journal over fake peer reviews. Post of April 20, 2017.

Schroter, S., Black, N., Evans, S., Godlee, F., Osorio, L., Smith, R. (2008): What errors do peer reviewers detect, and does training improve their ability to detect them? Journal of the Royal Society of Medicine 101(10):507–514. DOI: 10.1258/jrsm.2008.080062

Smith, R. (2021): Time to assume that health research is fraudulent until proved otherwise? BMJ Opinion, July 5, 2021.

Van Noorden, R. (2023): More than 10,000 research papers were retracted in 2023 – a new record. Nature 624:479–481. DOI: 10.1038/d41586-023-03974-8

von Samson-Himmelstjerna, G. (2012): Anthelmintic resistance in equine parasites – detection, potential clinical relevance and implications for control. Veterinary Parasitology 185(1):2–8. DOI: 10.1016/j.vetpar.2011.10.010

Wareham, K. J., Hyde, R. M., Grindlay, D., Brennan, M. L., Dean, R. S. (2017a): Sponsorship bias and quality of randomised controlled trials in veterinary medicine. BMC Veterinary Research 13:234. DOI: 10.1186/s12917-017-1146-9

Wareham, K. J., Hyde, R. M., Grindlay, D., Brennan, M. L., Dean, R. S. (2017b): Sample size and number of outcome measures of veterinary randomised controlled trials of pharmaceutical interventions funded by different sources, a cross-sectional study. BMC Veterinary Research 13:295. DOI: 10.1186/s12917-017-1207-0

Weese, J. S. (2026): Small sample sizes in clinical trials: a pragmatic approach to clinical research in veterinary medicine. Journal of Small Animal Practice. DOI: 10.1111/jsap.70063

Team Sanoanimal

Team Sanoanimal

We are an experienced team of therapists specializing in feed consultation and integrated therapies for horses. With extensive experience in treating metabolic issues, we focus on natural, species-appropriate feeding and proven naturopathic remedies to enhance your horse's health. Benefit from our expertise to ensure the well-being of your horse.

Further articles on this category

Advertisement
Search results are being compiled...