by Alex Roldan
The election fever has finally set in and the temperature is heating up. Undeniably, statistics are fueling the people’s fervor for the election next year. Poll surveys clearly dominate any election related discussion nowadays – either to assess the effectivity of the strategies of the political parties or to counter other statistical claims.
However, statistics done honestly can make a statement like no other; but done dishonestly become deeply deceptive because those who read them tend to believe that the numbers had been put together honestly. Poll results are potent weapons to change the public’s perception about the status of the political candidacy of an individual. If an aspirant finds his plans are in distress, all he has to do is “produce” statistics that show he is up there among the leading hopefuls!
Before the famous Dr. George H. Gallup invented random sampling to predict the result of an event like presidential election in USA, magazines like Literary Digest polled as many as millions of people. The 1936 Gallup Poll correctly predicted Franklin Delano Roosevelt’s victory over Alf Landon by polling at random only 5,000 persons while the Digest, which polled over 20 million of its readers, incorrectly predicted Landon’s victory. The reason for this was simple – readers of the Digest were not representative of the whole US population. Also, any questionnaire is only a sample (another level) of the possible question, and the answer to the question is no more than a sample (third level) of the respondents’ attitudes and experiences on each question.
But how do we distinguish good statistics from bad statistics? Let us remember that data can be manipulated to support any argument. If one wants to demonstrate that a certain candidate would likely win if elections were to be held at the time the survey is made, one merely changes the questions to be asked and the names of possible candidates. If the numbers ran up on one group don’t look good, pick another group. If the average is too low or too high, go for the median and arrange data to discard the high or low ends.
I would like to share with you some ways of how to distinguish good statistics from bad. I know that this is a difficult subject for some, but by trying to understand how polls are made helps one from having his legs pulled by these wily pollsters. One does not have to be a statistician to do this.
First, there are questions you should ask yourself before making any conclusion about a survey result, 1) is the question asked relevant? 2) did the data come from reliable sources? 3) are all the data reported, not just the best (or the worst)? 4) are the data presented in context? 5), have the data been interpreted correctly?
Second, common words that dominate the analysis of the figures in front of you are the words “average, median, and mean.” The word average has a very loose meaning. In English, the average may denote the mode (i.e. the value which occurs most frequently), the median (i.e. the value which is in the middle of the distribution, with 50% below and 50% over it), and the mean (arithmetical average). For normal (Gaussian) distribution mean, median, and mode fall at the same point. For a skewed or off center distribution, the mean could be quite a distance from the mode. Also the notion of standard deviation is closely related to the normal distribution. Those are the reasons why normal distribution is so often used without the necessary basis – it frees from careful consideration of the real meaning of the words one uses.
In dealing with averages, you should ask, Average of what? Who or what is included? What kind of Average is this? How accurate is the figure? Who says so, and how does he or she know? Meaning, don’t swallow the figures fed to you hook line and sinker!
Moreover, watch out for fallacious correlations because the cause-and effect nature of correlation is often only a matter of speculation – when “B” always follows “A,” it does not mean that A causes B. A co-variation may be real, but it may not be possible to be sure which of the variables is the cause and which the effect.
I am always tempted to go on about data manipulation, but I feel that it is no longer necessary – unless you want to be a full-fledged statistician, which I hate to be. But at this point, I will reveal to you my simplest and the most useful technique in trying to determine good from bad statistics.
All you need is a basic skill, the ability to determine the difference between “information” from “data.” Do not be confused, I am going to explain this further. For example, if you know a person who has money to spend on research to determine his chances if he decides to join politics. That is information, isn’t it? If one has that money already in his pocket, that’s data. Still confused?
Bad statistics is to extract data (money) or any favor from the information! If the purpose of the study is just to get the money or extract favor in the future, then, that is bad statistics. Unfortunately, many research papers nowadays show that much too often it is really the case. But who would you blame? Election is a season of plenty, and some of those with research skills are just trying to cash in on their expertise.
For comments, e-mail to roldanalex@yahoo.com
Subscribe
Login
0 Comments
Oldest



