Sunday, December 16, 2018

The Parachute Trial: useful caricature or just a joke?


Caricature studies" have been used successfully in the scientific field to make relevant methodological discussions more palatable. I like this approach and often use them as teaching tools, such as the strong correlation between chocolate consumption and Nobel Prizes as an example of confounding bias.

In 2003, a systematic review on efficacy of parachute use in patients who jumped from great heights was published in the British Medical Journal. The review indicated no randomized clinical trials for this intervention. It was a clever way of demonstrating that not everything needs experimental evidence. That article inspired us to create the terms "parachute paradigm" and "principle of extreme plausibility".

Yesterday, I received a plethora of enthusiastic messages about the latest clinical trial published in the British Medical Journal as part of the Christmas series: Parachute use to prevent death and major trauma when jumping from aircraft: randomized controlled trial.

In this trial, airplane passengers were invited to enter a study where they would jump from the plane to the ground, after being randomized to the use of parachute or non-parachute backpack as a control group. The primary outcome was death or severe trauma. Based on the premise that 99% of the control group would suffer the outcome, for a 99% power to detect a huge (and plausible) relative risk reduction of 95%, only 10 patients per group would be needed. This was done and, surprisingly, the study was negative: zero incidence of the primary outcome in both groups. However, only individuals who would jump from planes parked on the ground agreed to participate in the work.

Funny, but what is the implicit message of this study?

"Randomized trials might selectively enroll individuals with a lower perceived likelihood of benefit, thus diminishing the applicability of the results to clinical practice."

According to the authors, the new parachute study would be pointing to the problem that randomized clinical trials select samples less predisposed to the benefit of the intervention, a phenomenon that would promote false negative studies. The authors explain that it happens because patients who are more likely to benefit from therapy are less likely to agree to enter a study in which they may be randomized to non-treatment. This would make clinical trial samples less sensitive to benefit detection as a partial exclusion of patients with a greater chance of therapeutic success would take place.

Caricatures serve to accentuate true traits. However, if we were to characterize clinical trial samples (ideal world), they tend to be more predisposed to finding positive results in comparison with a real world target population. Therefore, this study is not a caricature of the real world clinical trial.

Thus, the present article should lose the caricature status and be considered just a funny joke, with no ability to anchor our mind towards a better scientific thinking.

As proof of concepts, clinical trials rely on the use of highly treatment-friendly samples, by applying restrictive inclusion and exclusion criteria. Differences between patients who accept and do not agree to enter the study are not sufficient to generate a sample less predisposed to treatment benefit than reality.

The "joke study" commits an unusual sample bias: it allows the inclusion of patients who do not need treatment. It would be as if a study aimed at testing thrombolysis allowed the inclusion of any chest pain, regardless of the electrocardiogram. Doctors who already believe in thrombolysis would see the electrocardiogram, thrombolyze ST-elevation patients, and release those who do not need thrombolysis to be randomized to drug or placebo. A joke without scientific value.

Caricature studies are useful when they anchor the mind of the community to a sharper criticism of the results of studies. However, in this case, the anchoring occurred in the opposite direction.

First, when we think of the scientific ecosystem, the biggest problem is false positive studies, mediated by several phenomena: confounding bias in observational studies, outcome reporting bias, conclusions skewed to positive finding  (spin) and, finally, citation bias that favor positive studies. Behind all this lies the innate predilection of the human mind for false statements, to the detriment of true denials.

Secondly, there is the problem of efficacy (ideal world) versus effectiveness (real world). Clinical trials aim to evaluate efficacy, which could be interpreted as the intrinsic potential of the intervention to offer clinical benefit: "Does the treatment have beneficial ownership?" Therefore clinical trials represent the ideal condition for the treatment to work. In the face of a positive clinical trial, we must always reflect whether this positivity will be reproduced in the real world, which constitutes effectiveness.

Of course there is the problem of false negative studies and it should also be a concern. But the bias suggested by the funny parachute study does not represent an important false-negative mechanism. The most prevalent mechanisms leading to false negatives are reduced statistical power, excessive crossover in the intention-to-treat analysis and inadequate intervention applicability.

My concern is that a reader of this funny study would take the following message home: if a promising study is negative, consider that clinical trials tend to include patients less likely to the benefit from the intervention. This message is wrong, as clinical trials tend to select samples more predisposed to the benefit. Of course, there are exceptions, but if we are to anchor our mind, it should be in the direction of the most prevalent phenomenon.

My prediction is that this study will come to be cited by legions of believers not satisfied with negative results from well-designed studies. Just as the seminal article of the parachute has been used inadequately as a justification for many treatments that have nothing to do with the parachute paradigm under the premise that "there is no evidence at all." A recent study by Vinay Prasad has shown that most interventions characterized as parachute paradigm by medical articles are not that, many have had clinical trials with negative results.

The great attention received by the parachute clinical trial is an example of how information sharing on social networks occurs. The main criterion for sharing is the interesting, unusual or amusing character, to the detriment of the veracity or usefulness of the information. In the appeal for novelty, fake news end up getting more attention than true news, as was recently demonstrated by a paper published in Science. Although the article we are discussing should not be framed as fake news, it is not a good caricature of the real world either.

The work in question is not a caricature of the ecosystem of randomized clinical trials. It is a mere joke with the potential to bias our minds to the inadequate idea that the heterogeneity between clinical trial samples and the target population of the treatment reduces the sensitivity of these studies to detect positive effects. In fact, the samples enrolled in clinical trials usually have a greater chance to detect positive results (sensitivity) than if the entire target population were included.


When the learning of science is approached in a fun way, it arouses great interest of the biomedical community. But we should always ask ourselves: what is the implicit message of the caricature? It is the first step to the critical appraisal of such “thought experiments”.

Saturday, November 3, 2018

The Bright Side of “Many Analysis, One Data Set” Paper



An elegant paper led by English researchers and recently published in the journal Advances in Methods and Practices in Psychological Science has enhanced scientific skepticism regarding ascertainment of statistical data analyses. Using exactly the same database, 29 independent research groups provided a priori data analysis plan to test the hypothesis that referees tend to give red cards more often to dark-skin-toned soccer players in comparison with light-skin-toned players. The analysis performed by 20 groups statistically confirmed the hypothesis, while 9 groups had non-significant statistical analysis.

Amidst of the scientific concern hype ignited by this paper, I have to confess that this time my interpretation leaned towards optimism. Considering the complexity of the problem analyzed, the observational nature of the data and the large variability of statistical methods chosen by the researchers, I found the results presented by different groups surprising similar. 

The authors described that odds ratio of dark-skin-toned players for getting red cards, in relation to light-skin-toned players, varied between 0.89 and 2.93. Although this interval appears to suggest high variation of results, by looking carefully at the forest plot depicted in the figure below, it becomes clear that most studies have similar odds ratios and confidence intervals. Actually, there were two outliers with odds ratio of 2.88 and 2.93 and extremely large confidence intervals. Something in those statistical analysis made these two studies very imprecise. On the other hand, the rest of studies had quite similar results.



Considering all 29 studies, we calculated an average odds ratio of 1.39, with 95% confidence interval between 1.22 and 1.55. If we exclude the two outliers, the average odds ratio is 1.28 (95% CI = 1.21 - 1.33, very precise). In reality, agreement among studies regarding both point estimate odds ratios and confidence intervals is quite good.

Furthermore, while 20 studies demonstrated a positive association between the dark-skin-toned players and odds to get a red card, no studies suggested the opposite result. The remaining 9 analysis basically did not reject the null hypothesis.

Assuming the true result is the one presented by most studies, none of the 9 discordant studies had made the most serious random error of claiming falsity (type I error). All 9 studies would have made the type II error, that is, they simply failed to reject the null hypothesis. Considering the association being explored is not strong (odds ratio < 2), it is only natural that some of the analyses lacked sufficient statistical power.

The problem presented to the researchers was quite complex. The observational nature of the data leads to potential confounding, along with concerns regarding independence of observations. Statistical analysis had to address heterogeneities between players according to skin-tone, referees predisposition to give red cards, relationship among players and referees, different soccer leagues, among other things. 

I may comply with a “half empty glass” interpretation of the study: choices for statistical approaches for complex epidemiological data vary substantially and this variation leads to a certain level of   disagreement among studies. On the other hand, I am more inclined to a “half full glass” interpretation: for a very complex problem, odds ratio estimation was surprisingly reproducible, most studies rejected the null hypothesis in the same direction and no studies suggested the opposite result. Moreover, if we take into consideration less complex statistical circumstances, such as the case for a typical well-designed large randomized controlled trial, the prospect may be quite good.

Friday, September 21, 2018

The Magical Transformation of a Secondary into a Primary Outcome



The reading of a scientific study should involve a domain beyond the scientific article, encompassing the ecosystem that involves the creation of the idea, definition of the protocol and acceptance of the results by the community. The reading of the work does not begin, nor does it end in the final article.

In a recent post, we provoked the reflection about the uncertain result of the SCOT-HEART clinical trial. That analysis was solely based on my reading of the article published in the NEJM. In the present article I will go further, beyond the final publication. 

In the journal club of my cardiology department, we use a peculiar methodology to read articles. One of these aspects is the orientation for our resident to systematically access clinicaltrial.gov and look for inconsistencies between the protocol defined a priori and what is in the published article. We are constantly evaluating the ecosystem prior to the article.

That was when our chief resident, Dr. JoĂŁo Menezes, came up with another surprise about the SCOT-HEART trial: the primary outcome reported in the NEJM publication was actually one of many secondary outcomes, exemplifying "the magical transformation of a secondary into a primary outcome".


The transformation

The scientific integrity of a study depends on a priori definition of data analysis. This method serves to avoid the multiplicity of tests that would increase the probability of type I error. In this context, it is essential to define the primary outcome of the study, which should guide the conclusion, instead of relying on the positive secondary outcomes results that may suffer from the multiple comparison problem.

The publication of SCOT-HEART in the NEJM clearly states "The primary endpoint was death from coronary heart disease or nonfatal myocardial infarction at 5 years."

"Our pre-specified primary long-term endpoint was the proportion of patients who died from coronary heart disease or had a nonfatal myocardial infarction at 5 years."

Let us now go to the ecosystem prior to the article. As we know, authors should record the protocol of any clinical trial prior to its execution and this is usually done at clinicaltrials.gov.

Upon checking the study protocol on clinicaltrials.gov, JoĂŁo realized that the primary outcome described in NEJM was not a true primary outcome! As in a magic spell, a prior secondary outcome was made primary in the description of the final article.

In fact, this study was originally designed to evaluate the proportion of patients who received a diagnosis of coronary disease, comparing tomography versus control strategy. This proportion was the primary outcome pre-defined by the study.

The secondary outcomes were divided into 5 domains (symptoms, diagnosis, additional investigations, treatment implemented, long-term clinical outcomes). In the domain of clinical outcomes, 9 secondary endpoints have been described, among which is the "cardiovascular death and nonfatal infarction" end-point, now described as primary in the NEJM article.

Look at the description of the secondary clinical outcomes as set out in clinicaltrials.gov and the Trials article describing the study design in 2012:

  1. Cardiovascular death or non-fatal Myocardial Infarction (MI) (ii) Cardiovascular death (iii) Non-fatal MI (iv) Cardiovascular death, non-fatal MI or non-fatal stroke (v) Non-fatal stroke (vi) All-cause death (vii) Coronary revascularisation; percutaneous coronary intervention or coronary artery bypass graft surgery (viii) Hospitalisation for chest pain including acute coronary syndromes and non-coronary chest pain (ix) Hospitalisation for cardiovascular disease including coronary artery disease, cerebrovascular disease and peripheral arterial disease.

To complicate matters further, clinical outcomes were pre-defined to be evaluated at 10-year follow-up and the article describes a 5-year follow-up. Thus, 5 years follow-up was not a priori definition. Strictly speaking, we are faced with a secondary outcome defined a posteriori (post-hoc analysis). This is not just semantics, because in the absence of a definition of when the outcome should be evaluated, we can test it year by year, hoping that chance presents us with a positive result at some point. At the moment the author is gifted by chance, he can prepare an abstract and submit to an important international congress. I'm not saying that's how it was done, I'm just showing what can be done with post-hoc endpoints.

In this way, we are facing a serious problem of multiple comparisons, which can be computed as follows:

Considering the 5% alpha, if the null hypothesis is true (tomography group = control group), the probability of a false positive result in a single primary outcome is 5%. However, we are making 9 secondary attempts to get a positive result. If each of these attempts has a 5% probability of a false positive result, the probability of a false positive result appearing in any of the outcomes is 1 - 0.95K, where K is the number of trials. Thus, the likelihood of any of these secondary outcomes being false-positive is 36%. Much higher than the 5% if we were analyzing a single primary outcome.


To aggravate, the statistical power of SCOT-HEART, after correction for the actual incidence of the outcome, is only 27%, as we mentioned in previous post. We then have two mechanisms of randomly manufacturing a false-positive: multiple endpoints tested and a study that lacks statistical power. In this way, the probability of false-positive becomes greater than 36%. Third, if we consider the risk of outcome bias (ascertained by electronic medical records, not adjudicated), SCOT-HEART is a random and systematic machine for generating false results.

This is one more explanation for the unlikely relative hazard reduction of 41% in the incidence of the combined outcome of infarction and cardiovascular death at 5 years of follow-up after coronary CT. As discussed in the previous article, the prevention of a clinical outcome by conducting an examination depends on three conditional probabilities (abnormal finding P x change in treatment P x ​​beneficial effect the changing treatment), different from the probability of benefit of a treatment that has only one component. To aggravate the gradient of changing treatment between the groups was only 4%.

In this way, it is too good to be true that conducting an examination promotes a benefit with the usual magnitude of good treatments, what usually ranges from 20% to 40%. Here we refer to the relative reduction because it describes the intrinsic "effect size" of a treatment, which does not vary with absolute risk. For example, the relative risk reduction of death for heart failure by enalapril or bata-blocker is 16% and 35%, respectively.


Two Scandals

It is scandalous for the authors to describe as primary outcome an outcome that was pre-defined as secondary. This shows the lack of scientific integrity behind the scenes of this work.

Perhaps more shocking is the acceptance of this article by the medical community, which seemed to commemorate the outcome of the work, featured prominently at the European Congress of Cardiology in Munich.

Problems of scientific integrity do not belong to a morally defective individual. Lack of scientific integrity stems from a faulty ecosystem, through research producers, editors and reviewers, and those who read the article without the necessary critical insight.


The Legion Bias

We might think it is very strange that thousands of cardiologists simultaneously attending the presentation of the study in ESC congress agreed with the doubtful result. Should the number of enthusiastic people be an evidence in favor of study truthfulness?

It is worth reminiscing to the observation of Swedish physician and statistician Hans Rosling, who became famous for his TED lectures, using dynamic statistical graphs to show how most people are wrong about important facts of life.

Rosling used to ask such questions to a legion of intellectuals: "How many children from low-income countries have basic education? 20%, 40%, or 60%?" The correct answer is 60%, but only 7% of intellectuals responded correctly. Most people chose 20%. Note that if we asked a monkey what the correct alternative would be, it would hit 33% of the time. Why do sapiens hit just 7%? The answer lies in our extreme bias. We tend to believe in the most significant result (most positive), whether we are talking about a risk factor or the beneficial effect of a treatment. Our mind has a tropism towards the highest possible contrast, making us choose the most extreme result.

It is a collective phenomenon, creating a legion of believers in the most significant result. The immense number of people thinking the same way, reinforces the belief of the legion participants. It's the legion bias.

The problem worsens when we are medical specialists, enthusiastic about our technological tools. This justifies the belief medical community deposited in small and biased studies of hypothermia post cardiac arrest and beta-blockers in non-cardiac surgery, which have become recommendations in guideline; or hormone replacement therapy from observational studies. The same is true for SCOT-HEART, which, when presented with glamor at the European Society of Cardiology Meeting, created its own legion of believers.


The Novelty, Positivism and Confirmation Biases

SCOT-HEART is the most recent study, so it appears as a novelty that promotes knowledge evolution. However, there was already another study published years earlier. This is the PROMISE study: a larger study (10,000 patients), a truly primary outcome defined a priori, with follow-up for evaluation of outcomes, adjudicated. That is, PROMISE is an immensely superior quality study to SCOT-HEART. And its result was negative.

Why, then, do we prefer to believe in the positive evidence of poor quality than in the negative evidence of good quality? Because our mind has a tropism for the positive (bias of positivism) and for the new (novelty bias). Then, we use the confirmation bias (we select positive evidence and disregard negative ones) to reinforce our belief.

By considering the cognitive biases of the biological mind, we do not need to be rude mentioning possible conflicts of interest that can also move the legions of believers.


The King Who Was Naked

It tells the story of Hans Christian Andersen (1937) that a very vain king ordered two tailors an unprecedented outfit, so original that no one had ever dressed the same. In the impossibility of realizing the king's desire, the tailors devised an imaginary costume, which they claimed to be invisible to the eyes of stupid people. The king himself, when he tried on his clothes, could not see it in the mirror, but pretended to see in order not to look stupid. In the same way, all people realized that the king was naked, but no one drew his attention for fear of being considered stupid. And so the king spent much of his reign naked, exposed to ridicule. The fear of looking stupid made people accept the unbelievable. In fact, many believed that they were seeing the clothes, because they wanted to believe that they were not stupid.

This story portrays the mechanism by which some myths persist in medicine.

One fine day, during an important parade in a public square, when a child saw the king passing by, he cried out: the king is naked! This child unmasked the charade created by the tailors, embarrassed the king, and especially the subjects who believed the lie or were ashamed to disagree.

Some interpret that it was the innocence of the child that allowed its observation. In fact, he was one of those half-malicious children. In this case, the difference between child and adult was the courage to acknowledge the truth and disagree with the legion of fanatics.

Let SCOT-HEART alert to the multiple biases that keep us from scientific integrity. 

Wednesday, August 29, 2018

Scientific Fake News: what is it exactly?



Improper information have always existed. By becoming a popular term, the expression "fake news" has alerted people and brought some useful skepticism, which is not a natural feature of the human mind. 

Evolutionary speaking, the human mind has evolved over 200.000 years of history to believe. In fact, evolutionary psychology claims that our unique ability to fantasise abstract phenomena was responsible for the species to prevail .  

As Francis Bacon once stated, "the human mind is more excited by affirmatives than negatives". And it was recently demonstrated in tweeter and published in Science: "fake news spread faster than true news".

Although we have evolved technologically and science is in the core of this evolution, the human mind did not have enough time to evolve from fantasy to skepticism. The last 500 hundred years were not enough to outrun 200.000 years of evolution. Biologically, we are believers.

The root of scientific thinking is skepticism. In science, we must have a method to overcome our predisposition to believe. This method is called the null hypothesis: we start by not believing and only change to the alternative of believe after strong evidence against chance or bias rejects the null. Being skeptical is tiresome and sometimes boring. 

This is in the center of a scientific problem: the lack of reproducibility, well described by Ioannidis in his popular PLOS One article: "most published research findings are false". And we believe them. 

The term "fake news" became popular two years ago and served as an alert for people, before becoming unpopular for political reasons. 

With a correct understanding its meaning, the term  "scientific fake news" helps against the problem of scientific reproducibility. But first, we must differentiate "scientific fake news" from "fake news". 

Fake news is created by a person or small group of people with common interest. Scientific fake news is created by a system who is defective: the creators are not alone, peer-reviewers, editors, societies and readers have to approve it and spread the message with enthusiasm. And they may do it with good intention.

Fake news has a creator who knows the news is fake. In scientific fake news, the creator believe in the message, a belief reinforced by his or her confirmation bias. 

Fake news has a creator with poor personal integrity. In scientific fake news, the creator suffers from scientific integrity, mediated biologically by cognitive bias. 

Fake news does not have empirical evidence, scientific fake news has experimental evidence that falsely suggests credibility.

Fake news is easily dismissed. Scientific fake news may take years to dismiss. It is responsible for the phenomenon of medical reversal, when improper information drives medical behaviour for years, only to be reverted after true and stronger evidence takes place. It was the case of medical therapies that were incorporated such as Xigris for sepses, hypothermia after cardiac arrest, beta-blockers for non-cardiac surgery and so on ... 

In his seminal article on medical reversal, Vinay Prasad wrote "we must raise the bar and before adopting medical technologies".

And the last difference: Donald Trump loves the term ''fake news", but has no ideia what "scientific fake news" mean. 

Well, it is not that scientific fake news is totally naive, there is also conflict of interest mediating it. But the main conflict comes from positivism bias, meaning every authors, editors or readers prefere positive studies over negative studies. 

Following description of human mind cognitive biases under uncertainty by Kahneman and Tversky, Richard Thaler came up with the solution to nudge human behaviour. Nudge means interventions to unconsciously change behaviour, with may be more effective than rational arguments. 

For example, to avoid people to cheat on tax returns, instead of explaining how important it is to pay taxes, a nudge would just say "most people fill their reports accurately". It was the most effective to improve behaviour in the UK. 

In the case of science, the expression fake news is so strong that may act as a nudge to scientific integrity. Yes, it may sound politically incorrect, but it is a disruptive nudge. Just speak about bias and chance has not been enough, as recently meant by Marcia Angell: "no longer possible to believe much of clinical research published".

Maybe we are not in a crises of scientific integrity. Actually, I think this type of discussions are becoming more frequent and we should be optimist.

But a nudge may accelerate the process: before reading any article, we should make a critical appraisal of our internal beliefs and ask ourselves: in this specific subject, am I specially vulnerable to believe in "scientific fake news"?

Monday, August 27, 2018

SCOT-HEART Trial: how to spot scientific fake news at a glance




It was just presented at the ESC Congress and simultaneously published in the NEJM a great example of scientific fake news, the SCOT-HEART Trial.

I use this example to show that the reading of an article starts before the traditional process. A pre-reading should bring us the critical spirit necessary for the reading process. During pre-reading we begin to develop a vision of the whole, as if we were looking at a city from the airplane window.

Then we'll land the plane and start reading to assess details.

The pre-reading of an article is composed of two questions: first, the hypothesis makes sense, should this study have been carried out? (pre-test probability of the idea = plausibility + previous studies); second, is the result too good to be true (effect size)?

In pre-reading process, we should avoid flooding the details head. We need only to identify the tested hypothesis and the main result. By reading just the conclusion of the article, we get this information which should be accompanied by a look at the line of results that presents the main numbers in order to get notion of effect size (it takes 30 seconds).

In the case of SCOT-HEART trial:

 "CTA in addition to standard care in patients with stable chest pain resulted in a significantly lower rate of death from coronary heart disease or nonfatal myocardial infarction at 5 years than standard care alone.

The 5-year rate of the primary end point was lower in the CTA group than in the standard care group (2.3% [48 patients] vs. 3.9% [81 patients]; hazard ratio, 0.59; 95% confidence interval [CI ], 0.41 to 0.84, P = 0.004)."

From these two sentences, we have noticed the hypothesis tested: the use of tomography in patients with stable thoracic pain reduces cardiovascular events. What is the pre-test probability of this idea?

There is some plausibility to the extent that anatomical information can modify the therapeutic behaviour of physicians and then modifies outcomes. Regarding previous evidence, the PROMISE study randomised 10,000 patients for tomography versus noninvasive evaluation and was negative for cardiovascular outcomes. The PROMISE control group is not exactly the same as SCOT-HEART, but indirectly the result of that study models negatively the pretest probability of the SCOT-HEART hypothesis. Therefore, I would say that the pre-test probability is low, but not zero, maintaining the justification for the study being performed.

Then comes the second question: is the effect size too good to be true? Note that the CT scan promoted a 41% relative hazard reduction. This magnitude of effect is typical of beneficial treatments. It is important to note that the effect size of a test will always be much less than that of a treatment, since in the first there are many more steps between intervention and outcome.

In the case of a clinical trial testing efficacy of a test, the following steps are necessary before the benefit occurs:

The examination is done on all patients - a portion of them has a result that may suggest to the physician to improve patient's treatment - in a sub-portion of these patients the physician actually enhances the treatment - a sub-sub-portion of patients benefits from treatment improve. Therefore, we should expect that the magnitude of the clinical effect of a test is much lower than that of a treatment.

In this way, we conclude that the SCOT-HEART result is too good to be true. It would be extraordinary for a test to promote such effect size. As Carl Sagan said, "extraordinary claims requires extraordinary evidence". Is the quality of this trial extraordinary?

Now let's read the article, looking for problems that justify such an unusual finding, 41% relative reduction of the hazard by performing an exam.

The first point that draws attention was the minimal difference in treatment modification promoted by the CT scan versus the control group. There was no difference in the revascularization procedure. Regarding preventive therapies such as statin or aspirin, the difference between the two groups was only 4% (19% versus 15%).

The number of patients in the CT scan group is 2,073 x 4% improvement in therapy = the CT scan group had an additional 83 patients with improved therapy in relation to control.

The number of events prevented in the CT group (relative to the control group) was 33.

Thus, drug enhancement of 83 patients prevented 33 clinical outcomes. If we were to evaluate the treatment that was performed at the end of the cascade, the NNT would be 2.5. Something unprecedented, that almost no real treatment is able to promote, nor a test.

This is a definitely false result.

The continuity of the reading will serve to understand the mechanisms that generated this false result.

"There were no trial-specific visits, and all follow-up information was obtained from data collected routinely by the Information and Statistics Division and the electronic Data Research and In- novation Service of the National Health Service (NHS) Scotland. These data include diagnostic codes from discharge records, which were classified according to the International Classification of Dis- eases, 10th Revision. There was no formal event adjudication, and end points were classified primarily on the basis of diagnostic codes."

So the outcomes were obtained through the electronic records review, through ICD and without adjucation by the authors. Second, the study was open and ascertainment bias can happen. For example, knowledge of a normal CT scan may influence the doctor who writes the ICD to interpret a symptom as innocent, while in another patient who is unaware of the anatomy, a symptom may prompt troponin measurement and subsequent diagnosis of nonfatal infarction. This is just a potential explanation, which serves as an example.

In fact, we are never able to open the black box of the exact mechanism that prevailed in generating a bias. However, it should be borne in mind that the combination of an open-label study with an inaccurate method of outcome measurement leads to a high risk of bias.

One of the techniques to explore the possibility of ascertainment bias is to compare the outcome of specific death (subject to scoring bias - subjectivity) with the result of death from any cause (immune to bias). Even though it is not a primary or statistically significant outcome, it is worth as exploratory analysis. It is interesting to note that the hazard ratio is 0.46 for cardiovascular death and 1.02 (null) for general death. In the absence of a substantial increase in non-cardiovascular death, this suggests that the study is especially subject to ascertainment bias for subjective outcomes.

In addition, the study presents a high risk of random error, since it is underpowered. In fact, the calculation of the sample was based on the premise of a 13% incidence of the outcome in the control group, but only 3.9% took place. By my calculation, it reduced a desired statistical power of 80% to 32%. As we know, small studies are more predisposed to false positive results because of their imprecision.

This imprecision not only increases the probability of type I error, but also incapacitates the study of measuring the size of the effect. That is, 41% relative reduction of the hazard presented a confidence interval ranging from 16% to 59%.

Finally, if we considered the information true, it would be worth an analysis of applicability. The hypothesis tested here is of a pragmatic nature. That is, an intervention is done at the beginning, and we expect the physician reacts in a way that benefits the patient. However, the protocol was designed to systematically influence physicians behaviour. 

"When there was evidence of nonobstructive (10 to 70%) cross-sectional luminal stenosis or obstructive coronary artery disease on the CTA, or when a patient had an ASSIGN score of 20 or higher, the attending clinician and primary care physician were prompted by the trial coordinating center to prescribe preventive therapies. "

This methodology reduces the external validity of the study, because we do not know if in the absence of this induction by the protocol of the study, doctors would act the same. If the benefit were true, in practice it would be of a smaller magnitude.

For studies of insufficient quality we should keep uncertainty in mind. But SCOT-HEART goes further: this is study is certainly false. A great example of fake scientific news.

Vitamin C for Sepsis: a philosophical-scientific point of view

The CITRIS-ALI trial was a negative trial recently published in the JAMA, which depicts a graphic figure with looks and numbers show a ...