Full Text Transcript (Pages 1–50 of 58)
(cid:1)(cid:2)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8) (cid:9)(cid:8) (cid:10)(cid:11)
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)
(cid:6)(cid:9)(cid:10)
(cid:3)(cid:4)(cid:11)(cid:3)(cid:4)(cid:12)(cid:12)(cid:8)(cid:2)(cid:9)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
LLLLLEEEEEAAAAARRRRRNNNNNIIIIINNNNNGGGGG OOOOOBBBBBJJJJJEEEEECCCCCTTTTTIIIIIVVVVVEEEEESSSSS
After reading this chapter a student will be able to understand–
(cid:1) The meaning of bivariate data and technique of preparation of bivariate distribution;
(cid:1) The concept of correlation between two variables and quantitative measurement of
correlation including the interpretation of positive, negative and zero correlation;
(cid:1) Concept of regression and its application in estimation of a variable from known set of
data.
1111122222.....11111 IIIIINNNNNTTTTTRRRRROOOOODDDDDUUUUUCCCCCTTTTTIIIIIOOOOONNNNN
In the previous chapter, we discussed many a statistical measure relating to Univariate distribution
i.e. distribution of one variable like height, weight, mark, profit, wages and so on. However,
there are situations that demand study of more than one variable simultaneously. A businessman
may be keen to know what amount of investment would yield a desired level of profit or a
student may want to know whether performing better in the selection test would enhance his or
her chance of doing well in the final examination. With a view to answering this series of questions,
we need study more than one variable at the same time. Correlation Analysis and Regression
Analysis are the two analysis that are made from a multivariate distribution i.e. a distribution of
more than one variable. In particular when there are two variables, say x and y, we study bivariate
distribution. We restrict our discussion to bivariate distribution only.
Correlation analysis, it may be noted, helps us to find an association or the lack of it between the
two variables x and y. Thus if x and y stand for profit and investment of a firm or the marks in
Statistics and Mathematics for a group of students, then we may be interested to know whether
x and y are associated or independent of each other. The extent or amount of correlation between
x and y is provided by different measures of Correlation namely Product Moment Correlation
Coefficient or Rank Correlation Coefficient or Coefficient of Concurrent Deviations. In Correlation
analysis, we must be careful about a cause and effect relation between the variables under
consideration because there may be situations where x and y are related due to the influence of
a third variable although no causal relationship exists between the two variables.
Regression analysis, on the other hand, is concerned with predicting the value of the dependent
variable corresponding to a known value of the independent variable on the assumption of a
mathematical relationship between the two variables and also an average relationship between
them.
1111122222.....22222 BBBBBIIIIIVVVVVAAAAARRRRRIIIIIAAAAATTTTTEEEEE DDDDDAAAAATTTTTAAAAA
When data are collected on two variables simultaneously, they are known as bivariate data
and the corresponding frequency distribution, derived from it, is known as Bivariate Frequency
Distribution. If x and y denote marks in Maths and Stats for a group of 30 students, then the
corresponding bivariate data would be (x, y) for i = 1, 2, …. 30 where (x , y ) denotes the
i i 1 1
marks in Maths and Stats for the student with serial number or Roll Number 1, (x , y ), that for
2 2
the student with Roll Number 2 and so on and lastly (x , y ) denotes the pair of marks for the
30 30
student bearing Roll Number 30.
(cid:10)(cid:11)(cid:12)(cid:11) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
As in the case of a Univariate Distribution, we need to construct the frequency distribution for
bivariate data. Such a distribution takes into account the classification in respect of both the
variables simultaneously. Usually, we make horizontal classification in respect of x and vertical
classification in respect of the other variable y. Such a distribution is known as Bivariate
Frequency Distribution or Joint Frequency Distribution or Two way Distribution of the two
variables x and y.
IIIIIlllllllllluuuuussssstttttrrrrraaaaatttttiiiiiooooonnnnn
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....11111Prepare a Bivariate Frequency table for the following data relating to the marks
in statistics (x) and Mathematics (y):
(15, 13), (1, 3), (2, 6), (8, 3), (15, 10), (3, 9), (13, 19),
(10, 11), (6, 4), (18, 14), (10, 19), (12, 8), (11, 14), (13, 16),
(17, 15), (18, 18), (11, 7), (10, 14), (14, 16), (16, 15), (7, 11),
(5, 1), (11, 15), (9, 4), (10, 15), (13, 12) (14, 17), (10, 11),
(6, 9), (13, 17), (16, 15), (6, 4), (4, 8), (8, 11), (9, 12),
(14, 11), (16, 15), (9, 10), (4, 6), (5, 7), (3, 11), (4, 16),
(5, 8), (6, 9), (7, 12), (15, 6), (18, 11), (18, 19), (17, 16)
(10, 14),
Take mutually exclusive classification for both the variables, the first class interval being 0-4 for
both.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
From the given data, we find that
Range for x = 19–1 = 18
Range for y = 19–1 = 18
We take the class intervals 0-4, 4-8, 8-12, 12-16, 16-20 for both the variables. Since the first pair
of marks is (15, 13) and 15 belongs to the fourth class interval (12-16) for x and 13 belongs to
the fourth class interval for y, we put a stroke in the (4, 4)-th cell. We carry on giving tally
marks till the list is exhausted.
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
TTTTTaaaaabbbbbllllleeeee 1111122222.....11111
Bivariate Frequency Distribution of Marks of Statistics and Mathematics.
MARKS IN MATHS
Y 0-4 4-8 8-12 12-16 16-20 Total
X
0–4 I (1) I (1) II (2) 4
MARKS 4–8 I (1) IIII (4) IIII (5) I (1) I (1) 12
IN STATS
8–12 I (1) II (2) IIII (4) IIII I (6) I (1) 14
12–16 I (1) III (3) II (2) IIII (5) 11
16–20 I (1) IIII (5) III (3) 9
Total 3 8 15 14 10 50
We note, from the above table, that some of the cell frequencies (f ) are zero. Starting from the
ij
above Bivariate Frequency Distribution, we can obtain two types of univariate distributions
which are known as:
(a) Marginal distribution.
(b) Conditional distribution.
If we consider the distribution of stat marks along with the marginal totals presented in the
last column of Table 12-1, we get the marginal distribution of marks of Statistics. Similarly, we
can obtain one more marginal distribution of Mathematics marks. The following table shows
the marginal distribution of marks of Statistics.
TTTTTaaaaabbbbbllllleeeee 1111122222.....22222
Marginal Distribution of Marks of Statistics
Marks No. of Students
0-4 4
4-8 12
8-12 14
12-16 11
16-20 9
Total 50
We can find the mean and standard deviation of marks of Statistics from Table 12.2. They
would be known as marginal mean and marginal SD of stats marks. Similarly, we can obtain
the marginal mean and marginal SD of Maths marks. Any other statistical measure in respect
of x or y can be computed in a similar manner.
(cid:10)(cid:11)(cid:12)(cid:21) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
If we want to study the distribution of Stat Marks for a particular group of students, say for
those students who got marks between 8 to 12 in Maths, we come across another univariate
distribution known as conditional distribution.
TTTTTaaaaabbbbbllllleeeee 1111122222.....33333
Conditional Distribution of Marks in Statistics for Students
having Mathematics Marks between 8 to 12
Marks No. of Students
0-4 2
4-8 5
8-12 4
12-16 3
16-20 1
Total 15
We may obtain the mean and SD from the above table. They would be known as conditional
mean and conditional SD of marks of Statistics. The same result holds for marks of Mathematics.
In particular, if there are m classification for x and n classifications for y, then there would be
altogether (m + n) conditional distribution.
1111122222.....33333 CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN AAAAANNNNNAAAAALLLLLYYYYYSSSSSIIIIISSSSS
While studying two variables at the same time, if it is found that the change in one variable is
reciprocated by a corresponding change in the other variable either directly or inversely, then
the two variables are known to be associated or correlated. Otherwise, the two variables are
known to be dissociated or uncorrelated or independent. There are two types of correlation.
(i) Positive correlation
(ii) Negative correlation
If two variables move in the same direction i.e. an increase (or decrease) on the part of one
variable introduces an increase (or decrease) on the part of the other variable, then the two
variables are known to be positively correlated. As for example, height and weight yield and
rainfall, profit and investment etc. are positively correlated.
On the other hand, if the two variables move in the opposite directions i.e. an increase (or a
decrease) on the part of one variable result a decrease (or an increase) on the part of the other
variable, then the two variables are known to have a negative correlation. The price and demand
of an item, the profits of Insurance Company and the number of claims it has to meet etc. are
examples of variables having a negative correlation.
The two variables are known to be uncorrelated if the movement on the part of one variable
does not produce any movement of the other variable in a particular direction. As for example,
Shoe-size and intelligence are uncorrelated.
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:22)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
1111122222.....44444 MMMMMEEEEEAAAAASSSSSUUUUURRRRREEEEESSSSS OOOOOFFFFF CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN
We consider the following measures of correlation:
(a) Scatter diagram
(b) Karl Pearson’s Product moment correlation coefficient
(c) Spearman’s rank correlation co-efficient
(d) Co-efficient of concurrent deviations
(((((aaaaa))))) SSSSSCCCCCAAAAATTTTTTTTTTEEEEERRRRR DDDDDIIIIIAAAAAGGGGGRRRRRAAAAAMMMMM
This is a simple diagrammatic method to establish correlation between a pair of variables.
Unlike product moment correlation co-efficient, which can measure correlation only when the
variables are having a linear relationship, scatter diagram can be applied for any type of
correlation – linear as well as non-linear i.e. curvilinear. Scatter diagram can distinguish between
different types of correlation although it fails to measure the extent of relationship between the
variables.
Each data point, which in this case a pair of values (x, y) is represented by a point in the
i i
rectangular axis of ordinates. The totality of all the plotted points forms the scatter diagram.
The pattern of the plotted points reveals the nature of correlation. In case of a positive correlation,
the plotted points lie from lower left corner to upper right corner, in case of a negative correlation
the plotted points concentrate from upper left to lower right and in case of zero correlation,
the plotted points would be equally distributed without depicting any particular pattern. The
following figures show different types of correlation and the one to one correspondence between
scatter diagram and product moment correlation coefficient.
Y
Y
O X O X
FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....11111 FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....22222
SSSSShhhhhooooowwwwwiiiiinnnnnggggg PPPPPooooosssssiiiiitttttiiiiivvvvveeeee CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn SSSSShhhhhooooowwwwwiiiiinnnnnggggg pppppeeeeerrrrrfffffeeeeecccccttttt
(((((00000 <<<<< rrrrr <<<<<11111))))) (((((rrrrr ===== 11111)))))
(cid:10)(cid:11)(cid:12)(cid:23) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
Y
Y
O X O X
FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....33333 FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....44444
SSSSShhhhhooooowwwwwiiiiinnnnnggggg NNNNNeeeeegggggaaaaatttttiiiiivvvvveeeee CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn SSSSShhhhhooooowwwwwiiiiinnnnnggggg pppppeeeeerrrrrfffffeeeeecccccttttt NNNNNeeeeegggggaaaaatttttiiiiivvvvveeeee
CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn
(((((–––––11111 <<<<< rrrrr <<<<<00000))))) (((((rrrrr ===== –––––11111)))))
Y
Y
O X O X
FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....55555 FFFFFIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....66666
SSSSShhhhhooooowwwwwiiiiinnnnnggggg NNNNNooooo CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn SSSSShhhhhooooowwwwwiiiiinnnnnggggg CCCCCuuuuurrrrrvvvvviiiiillllliiiiinnnnneeeeeaaaaarrrrr
CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn
(((((rrrrr ===== 00000))))) (((((rrrrr ===== 00000)))))
(((((bbbbb))))) KKKKKAAAAARRRRRLLLLL PPPPPEEEEEAAAAARRRRRSSSSSOOOOONNNNN’’’’’SSSSS PPPPPRRRRROOOOODDDDDUUUUUCCCCCTTTTT MMMMMOOOOOMMMMMEEEEENNNNNTTTTT CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN CCCCCOOOOOEEEEEFFFFFFFFFFIIIIICCCCCIIIIIEEEEENNNNNTTTTT
This is by for the best method for finding correlation between two variables provided the
relationship between the two variables in linear. Pearson’s correlation coefficient may be defined
as the ratio of covariance between the two variables to the product of the standard deviations
of the two variables. If the two variables are denoted by x and y and if the corresponding
bivariate data are (x, y) for i = 1, 2, 3, ….., n, then the coefficient of correlation between x and
i i
y, due to Karl Pearson, in given by :
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:24)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
( )
Cov x,y
r = r = .........................................................................(12.1)
xy ×
S S
x y
Where
∑ (x –x)(y –y) ∑xy
cov (x, y) = i i = i i –xy.............(12.2)
n n
2
S =
∑ (x
i
–x)
=
∑x
i
2
–x
2
..................................................(12.3)
x
n n
2
( )
and S = ∑ y i – y = ∑y i 2 – y2 .........................................(12.4)
y n n
A single formula for computing correlation coefficient is given by
n∑x y –∑x ×∑y
r= i i i i
( )2
n∑x2 – ∑x n∑y2 –( ∑ y )2 .............................................(12.5)
i i i
i
In case of a bivariate frequency distribution, we have
∑x yf
i i ij
i,j
Cov(x,y)= –x×y…………………………………...………(12.6)
N
∑f x2
io i
S = i – x2 .........................................................................(12.7)
x
N
∑f y2
oj j
and S = j – y2 ........................................................................(12.8)
y
N
Where x = Mid-value of the ith class interval of x
i
(cid:10)(cid:11)(cid:12)(cid:25) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
y = Mid-value of the jth class interval of y
j
f = Marginal frequency of x
io
f = Marginal frequency of y
oj
f = frequency of the (i, j)th cell
ij
∑f ∑f ∑f
N = ij = io = oj = Total frequency............... (12.9)
i,j i j
PPPPPRRRRROOOOOPPPPPEEEEERRRRRTTTTTIIIIIEEEEESSSSS OOOOOFFFFF CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN CCCCCOOOOOEEEEEFFFFFFFFFFIIIIICCCCCIIIIIEEEEENNNNNTTTTT
(i) TTTTThhhhheeeee CCCCCoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnnttttt ooooofffff CCCCCooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn iiiiisssss aaaaa uuuuunnnnniiiiittttt-----fffffrrrrreeeeeeeeee mmmmmeeeeeaaaaasssssuuuuurrrrreeeee.....
This means that if x denotes height of a group of students expressed in cm and y denotes
their weight expressed in kg, then the correlation coefficient between height and weight
would be free from any unit.
(ii) TTTTThhhhheeeee cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnnttttt ooooofffff cccccooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn rrrrreeeeemmmmmaaaaaiiiiinnnnnsssss iiiiinnnnnvvvvvaaaaarrrrriiiiiaaaaannnnnttttt uuuuunnnnndddddeeeeerrrrr aaaaa ccccchhhhhaaaaannnnngggggeeeee ooooofffff ooooorrrrriiiiigggggiiiiinnnnn aaaaannnnnddddd/////ooooorrrrr ssssscccccaaaaallllleeeee
ooooofffff ttttthhhhheeeee vvvvvaaaaarrrrriiiiiaaaaabbbbbllllleeeeesssss uuuuunnnnndddddeeeeerrrrr cccccooooonnnnnsssssiiiiidddddeeeeerrrrraaaaatttttiiiiiooooonnnnn.....
This property states that if the original pair of variables x and y is changed to a new pair of
variables u and v by effecting a change of origin and scale for both x and y i.e.
x−a y−c
u= andv=
b d
Where a and c are the origins of x and y and b and d are the respective scales and then we have
bd
r = r
xy uv ....................................................................(12.10)
b d
r and r being the coefficient of correlation between x and y and u and v respectively, (12.10)
xy uv
established, numerically, the two correlation coefficients remain equal and they would have
opposite signs only when b and d, the two scales, differ in sign.
(iii) TTTTThhhhheeeee cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnnttttt ooooofffff cccccooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn aaaaalllllwwwwwaaaaayyyyysssss llllliiiiieeeeesssss bbbbbeeeeetttttwwwwweeeeeeeeeennnnn –––––11111 aaaaannnnnddddd 11111,,,,, iiiiinnnnncccccllllluuuuudddddiiiiinnnnnggggg bbbbbooooottttthhhhh ttttthhhhheeeee llllliiiiimmmmmiiiiitttttiiiiinnnnnggggg
vvvvvaaaaallllluuuuueeeeesssss iiiii.....eeeee.....
–1 ≤ r ≤ 1 ………………… .............................................(12.11)
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....22222 Compute the correlation coefficient between x and y from the following data n
= 10, ∑xy = 220, ∑x2 = 200, ∑y2 = 262
∑x = 40 and ∑y = 50
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:26)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
From the given data, we have applying (12.5),
n∑xy–∑x×∑y
r = n∑x 2 –( ∑x) 2 × n∑y 2 –( ∑y) 2
10×220−40×50
=
10×200−(40)2× 10×262−(50)2
2200−2000
=
2000−1600× 2620−2500
200
=
20×10.9545
= 0.91
Thus there is a good amount of positive correlation between the two variables x and y.
AAAAAlllllttttteeeeerrrrrnnnnnaaaaattttteeeeelllllyyyyy
∑x 40
As given, x= = =4
n 10
∑y 50
y= = =5
n 10
∑xy
Cov (x, y) = −x.y
n
220
= −4.5=2
10
∑x2
S =
−(x)2
x n
200
=
−42 =2
10
(cid:10)(cid:11)(cid:12)(cid:10)(cid:27) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
∑y2
S =
i −y2
y n
262
=
−52
10
= 26.20−25=1.0954
Thus applying formula (12.1), we get
cov(x,y)
r =
S .S
x y
2
= = 0.91
2×1.0954
As before, we draw the same conclusion.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....33333 Find product moment correlation coefficient from the following information:
X : 2 3 5 5 6 8
Y : 9 8 8 6 5 3
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
In order to find the covariance and the two standard deviation, we prepare the following
table:
TTTTTaaaaabbbbbllllleeeee 1111122222.....33333
Computation of Correlation Coefficient
x y xy x 2 y2
i i i i i i
(1) (2) (3)= (1) x (2) (4)= (1)2 (5)= (2)2
2 9 18 4 81
3 8 24 9 64
5 8 40 25 64
5 6 30 25 36
6 5 30 36 25
8 3 24 64 9
29 39 166 163 279
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:10)(cid:10)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
We have
29 39
x= = 4.8333 y= =6.50
6 6
∑x y
cov (x, y) = i i −xy
n
= 166/6 – 4.8333 × 6.50 = –3.7498
Σx2
= i −(x)2
n
163
=
−(4.8333)2
6
= 27.1667–23.3608=1.95
∑y2
S = i −(y)2
y n
279
=
−(6.50)2
6
= 46.50−42.25=2.0616
Thus the correlation coefficient between x and y in given by
cov(x,y)
r =
S ×s
x y
–3.7498
=
1.9509×2.0616
= –0.93
We find a high degree of negative correlation between x and y. Also, we could have applied
formula (12.5) as we have done for the first problem of computing correlation coefficient.
Sometimes, a change of origin reduces the computational labor to a great extent. This we are
going to do in the next problem.
(cid:10)(cid:11)(cid:12)(cid:10)(cid:11) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....44444 The following data relate to the test scores obtained by eight salesmen in an
aptitude test and their daily sales in thousands of rupees:
Salesman : 1 2 3 4 5 6 7 8
scores : 60 55 62 56 62 64 70 54
Sales : 31 28 26 24 30 35 28 24
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Let the scores and sales be denoted by x and y respectively. We take a, origin of x as the average
of the two extreme values i.e. 54 and 70. Hence a = 62 similarly, the origin of y is taken
24+ 35
as the ≅ 30
2
TTTTTaaaaabbbbbllllleeeee 1111122222.....44444
Computation of Correlation Coefficient Between Test Scores and Sales.
Scores Sales in u v uv u2 v2
i i i i i i
(x) Rs. 1000 = x – 62 = y – 30
i i i
(1) (y)
i
(2) (3) (4) (5)=(3)x(4) (6)=(3) 2 (7)=(4) 2
60 31 –2 1 –2 4 1
55 28 –7 –2 14 49 4
62 26 0 –4 0 0 16
56 24 –6 –6 36 36 36
62 30 0 0 0 0 0
64 35 2 5 10 4 25
70 28 8 –2 –16 64 4
54 24 –8 –6 48 64 36
Total — –13 –14 90 221 122
Since correlation coefficient remains unchanged due to change of origin, we have
n∑u v −∑u ×∑v
i i i i
r = r xy = r uv = n∑u 2−
(
∑u
)2
× n∑v 2−
(
∑v
)2
i i i i
8×90−(−13)×(−14)
=
8×221−(−13)2× 8×122−(−14)2
538
=
1768 − 169 × 976 − 196
= 0.48
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:10)(cid:20)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
In some cases, there may be some confusion about selecting the pair of variables for which
correlation is wanted. This is explained in the following problem.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....55555 Examine whether there is any correlation between age and blindness on the
basis of the following data:
Age in years : 0-10 10-20 20-30 30-40 40-50 50-60 60-70 70-80
No. of Persons
(in thousands) : 90 120 140 100 80 60 40 20
No. of blind Persons :10 15 18 20 15 12 10 06
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Let us denote the mid-value of age in years as x and the no. of blind persons per lakh as y. Then
as before, we compute correlation coefficient between x and y.
TTTTTaaaaabbbbbllllleeeee 1111122222.....55555
Computation of correlation between age and blindness
Age in Mid-value No. of No. of No. of xy x2 y2
years x Persons blind blind per (2)×(5) (2)2 (5)2
(1) (2) (‘000) B lakh (6) (7) (8)
P (4) y=B/P × 1 lakh
(3) (5)
0-10 5 90 10 11 55 25 121
10-20 15 120 15 12 180 225 144
20-30 25 140 18 13 325 625 169
30-40 35 100 20 20 700 1225 400
40-50 45 80 15 19 855 2025 361
50-60 55 60 12 20 1100 3025 400
60-70 65 40 10 25 1625 4225 625
70-80 75 20 6 30 2250 5625 900
Total 320 — — 150 7090 17000 3120
(cid:10)(cid:11)(cid:12)(cid:10)(cid:21) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
The correlation coefficient between age and blindness is given by
n∑xy−∑x.∑y
r =
n∑x2 −(∑x)2 × n∑y2 −(∑y)2
8.7090−320.150
=
8.17000−(320)2 × 8.3120−(150)2
8720
=
183.3030.49.5984
= 0.96
Which exhibits a very high degree of positive correlation between age and blindness.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....66666 Coefficient of correlation between x and y for 20 items is 0.4. The AM’s and SD’s
of x and y are known to be 12 and 15 and 3 and 4 respectively. Later on, it was found that the
pair (20, 15) was wrongly taken as (15, 20). Find the correct value of the correlation coefficient.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
We are given that n = 20 and the original r = 0.4, x = 12, y = 15, S = 3 and S = 4
x y
cov(x,y) cov(x,y)
=0.4=
r =
S ×S 3×4
x y
= Cov (x, y) = 4.8
∑xy
= −xy=4.8
n
∑xy
= −12×15=4.8
20
= ∑xy = 3696
Hence, corrected = 3696 – 20 × 15 + 15 × 20 = 3696
Also, S 2 = 9
x
= (∑x2/ 20) – 122 = 9
= ∑x2 = 3060
Similarly, S 2 = 16
y
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:10)(cid:22)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
2
∑y 2
= −15 =16
20
= ∑ y2 = 4820
Thus corrected ∑x = n x – wrong x value + correct x value.
= 20 × 12 – 15 + 20
= 245
Similarly corrected∑y = 20 × 15 – 20 + 15 = 295
Corrected ∑x2 = 3060 – 152 + 202 = 3235
Corrected ∑y2 = 4820 – 202 + 152 = 4645
Thus corrected value of the correlation coefficient by applying formula (12.5)
20.3696−245.295
=
20.3235−(245)2 × 20.4645−(295)2
73920−72275
=
68.3740×76.6480
= 0.31
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....77777 Compute the coefficient of correlation between marks in Stats and Maths for the
bivariate frequency distribution shown in table 12.1
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
For the save of computational advantage, we effect a change of origin and scale for both the
variable x and y.
x −a x −10
Define u = i = i
i b 4
y −c y −10
And v = i = i
j d 4
Where x and y denote respectively the mid-values of the x-class interval and y-class interval
i j
respectively. The following table shows the necessary calculation on the right top corner of
each cell, the product of the cell frequency, corresponding u value and the respective v value
has been shown. They add up in a particular row or column to provide the value of f uv for
ij i j
that particular row or column.
TTTTTaaaaabbbbbllllleeeee 1111122222.....66666
Computation of Correlation Coefficient Between Marks of Maths and Stats
(cid:10)(cid:11)(cid:12)(cid:10)(cid:23) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
Class Interval 0-4 4-8 8-12 12-16 16-20
Mid-value 2 6 10 14 18
Class Mid V f f u f u2 f uv
j io io i io i ij i j
Interval-value u –2 –1 0 1 2
i
0-4 2 –2 1 4 1 2 2 0 4 –8 16 6
4-8 6 –1 2 4 4 4 5 0 1 –1 1 –2 13 –13 13 5
8-12 10 0 2 0 4 0 6 0 1 0 13 0 0 0
12-16 14 1 1 –1 3 0 2 2 5 10 11 11 11 11
16-20 18 2 1 0 5 10 3 12 9 18 36 22
foj 3 8 15 14 10 50 5 76 44
f v –6 –8 0 14 20 20
oj j
f v2 12 8 0 14 40 74
oj j
f uv 8 5 0 11 20 44 CHECK
ij i j
A single formula for computing correlation coefficient from bivariate frequency distribution is
given by
N∑f u v –∑f u ×∑f v
ij i j io i oj j
i,j
r = N∑f u2 –( ∑f u )2 ×∑f v2 – ( ∑f v )2 ...........................( 12. 10)
io i io i oj j oj j
50×44−8×20
=
50×76−82 50×74−202
2040
=
61.1228×57.4456
= 0.58
The value of r shown a good amount of positive correlation between the marks in Statistics
and Mathematics on the basis of the given data.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....88888 Given that the correlation coefficient between x and y is 0.8, write down the
correlation coefficient between u and v where
(i) 2u + 3x + 4 = 0 and 4v + 16x + 11 = 0
(ii) 2u – 3x + 4 = 0 and 4v + 16x + 11 = 0
(iii) 2u – 3x + 4 = 0 and 4v – 16x + 11 = 0
(iv) 2u + 3x + 4 = 0 and 4v – 16x + 11 = 0
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:10)(cid:24)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Using (12.10), we find that
bd
r = r
xy b d uv
i.e. r = r if b and d are of same sign and r = –r when b and d are of opposite signs, b and
xy uv uv xy
d being the scales of x and y respectively. In (i), u = (–2) + (-3/2) x and v = (–11/4) + (–4)y.
Since b = –3/2 and d = –4 are of same sign, the correlation coefficient between u and v would
be the same as that between x and y i.e. r = 0.8 =r
xy uv
In (ii), u = (–2) + (3/2)x and v = (–11/4) + (–4)y Hence b = 3/2 and d = –4 are of opposite signs
and we have r = –r = –0.8
uv xy
Proceeding in a similar manner, we have r = 0.8 and – 0.8 in (iii) and (iv).
uv
(((((ccccc))))) SSSSSPPPPPEEEEEAAAAARRRRRMMMMMAAAAANNNNN’’’’’SSSSS RRRRRAAAAANNNNNKKKKK CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN CCCCCOOOOOEEEEEFFFFFFFFFFIIIIICCCCCIIIIIEEEEENNNNNTTTTT
When we need finding correlation between two qualitative characteristics, say, beauty and
intelligence, we take recourse to using rank correlation coefficient. Rank correlation can also
be applied to find the level of agreement (or disagreement) between two judges so far as assessing
a qualitative characteristic is concerned. As compared to product moment correlation coefficient,
rank correlation coefficient is easier to compute, it can also be advocated to get a first hand
impression about the correlation between a pair of variables.
Spearman’s rank correlation coefficient is given by
6 ∑ d2
r = 1 − i ........................................... (12.11)
R
n(n2−
1)
Where r denotes rank correlation coefficient and it lies between –1 and 1.
R
d = x – y represents the difference in ranks for the i-th individual and n denotes the no. of
i i i
individuals.
In case u individuals receive the same rank, we describe it as a tied rank of length u. In case of
a tied rank, formula (12.11) is changed to
(cid:10)(cid:11)(cid:12)(cid:10)(cid:25) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
( )
6
⎡
⎢∑d2+∑
tj3−
tj
⎤
⎥
⎢ ⎣ i i j 12 ⎥ ⎦
r = − ................................................... (12.12)
1 ( )
R
n n2 − 1
In this formula, t represents the jth tie length and the summation
∑(t
j
3–t
j
)
extends over the
j j
lengths of all the ties for both the series.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....99999 compute the coefficient of rank correlation between sales and advertisement
expressed in thousands of rupees from the following data:
Sales : 90 85 68 75 82 80 95 70
Advertisement : 7 6 2 3 4 5 8 1
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Let the rank given to sales be denoted by x and rank of advertisement be denoted by y. We note
that since the highest sales as given in the data, is 95, it is to be given rank 1, the second highest
sales 90 is to be given rank 2 and finally rank 8 goes to the lowest sales, namely 68. We have
given rank to the other variable advertisement in a similar manner. Since there are no ties, we
apply formula (12.11).
TTTTTaaaaabbbbbllllleeeee 1111122222.....77777
Computation of Rank correlation between Sales and Advertisement.
Sales Advertisement Rank for Rank for d = x – y d2
i i i i
Sales (x) Advertisement
i
(y)
i
90 7 2 2 0 0
85 6 3 3 0 0
68 2 8 7 1 1
75 3 6 6 0 0
82 4 4 5 –1 1
80 5 5 4 1 1
95 8 1 1 0 0
70 1 7 8 –1 1
Total — — — 0 4
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:10)(cid:26)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
Since n = 8 and ∑d2 = 4, applying formula (12.11), we get.
i
6∑d2
r = 1− i
R n(n2−1)
6×4
1−
=
8(82 −1)
= 1–0.0476
= 0.95
The high positive value of the rank correlation coefficient indicates that there is a very good
amount of agreement between sales and advertisement.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111100000 Compute rank correlation from the following data relating to ranks given by
two judges in a contest:
Serial No. of Candidate : 1 2 3 4 5 6 7 8 9 10
Rank by Judge A : 10 5 6 1 2 3 4 7 9 8
Rank by Judge B : 5 6 9 2 8 7 3 4 10 1
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
We directly apply formula (12.11) as ranks are already given.
TTTTTaaaaabbbbbllllleeeee 1111122222.....88888
Computation of Rank Correlation Coefficient between the ranks given by 2 Judges
Serial No. Rank by A (x
i
) Rank by B (y
i
) d
i
= x
i
– y
i
d2
i
1 10 5 5 25
2 5 6 –1 1
3 6 9 –3 9
4 1 2 –1 1
5 2 8 –6 36
6 3 7 –4 16
7 4 3 1 1
8 7 4 3 9
9 8 10 –2 4
10 9 1 8 64
Total — — 0 166
(cid:10)(cid:11)(cid:12)(cid:11)(cid:27) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
The rank correlation coefficient is given by
6∑d2
r = 1− i
R n(n2–1)
6×166
1−
=
10(102 −1)
= –0.006
The very low value (almost 0) indicates that there is hardly any agreement between the ranks
given by the two Judges in the contest.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222 .....1111111111 Compute the coefficient of rank correlation between Eco. marks and stats.
Marks as given below:
Eco Marks : 80 56 50 48 50 62 60
Stats Marks : 90 75 75 65 65 50 65
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
This is a case of tied ranks as more than one student share the same mark both for Eco and
stats. For Eco. the student receiving 80 marks gets rank 1 one getting 62 marks receives rank 2,
the student with 60 receives rank 3, student with 56 marks gets rank 4 and since there are two
students, each getting 50 marks, each would be receiving a common rank, the average of the
5+6
next two ranks 5 and 6 i.e. i.e. 5.50 and lastly the last rank..
2
7 goes to the student getting the lowest Eco marks. In a similar manner, we award ranks to the
students with stats marks.
TTTTTaaaaabbbbbllllleeeee 1111122222.....99999
Computation of Rank Correlation Between Eco Marks and Stats Marks with Tied Marks
Eco Mark Stats Mark Rank for Eco Rank for d
i
= x
i
– y
i
d2
i
(x) (y) Stats
i i
80 90 1 1 0 0
56 75 4 2.50 1.50 2.25
50 75 5.50 2.50 3 9
48 65 7 5 2 4
50 65 5.50 5 0.50 0.25
62 50 2 7 –5 25
60 65 3 5 –2 4
Total — — — 0 44.50
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:11)(cid:10)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
For Eco mark there is one tie of length 2 and for stats mark, there are two ties of lengths 2 and
3 respectively.
( )
( ) ( ) ( )
∑ t3−t 23 −2 + 23 −2 + 33 −3
j j
Thus = =3
12 12
( )
6
⎡
⎢∑d2+∑
tj3−
tj
⎤
⎥
⎢ ⎣ i i j 12 ⎥ ⎦
Thus r = −
1 ( )
R
n n2 − 1
6×(44.50+3)
1−
=
7(72−1)
= 0.15
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111122222 For a group of 8 students, the sum of squares of differences in ranks for Maths
and stats marks was found to be 50 what is the value of rank correlation coefficient?
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
As given n = 8 and ∑d2 = 50. Hence the rank correlation coefficient between marks in Maths
i
and stats is given by
6∑d2
1− i
r R = n ( n2 −1 )
6×50
1−
=
8(82 −1)
= 0.40
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111133333 For a number of towns, the coefficient of rank correlation between the people
living below the poverty line and increase of population is 0.50. If the sum of squares of the
differences in ranks awarded to these factors is 82.50, find the number of towns.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
As given r = 0.50, ∑d2 = 82.50.
R i
6∑d2
1− i
Thus r R = n ( n2 −1 )
(cid:10)(cid:11)(cid:12)(cid:11)(cid:11) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
6×82.50
1−
0.50 = n ( n2−1 )
= n (n2 – 1) = 990
= n (n2 – 1) = 10(102 – 1)
∴ n = 10 as n must be a positive integer.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111144444 While computing rank correlation coefficient between profits and investment
for 10 years of a firm, the difference in rank for a year was taken as 7 instead of 5 by mistake
and the value of rank correlation coefficient was computed as 0.80. What would be the correct
value of rank correlation coefficient after rectifying the mistake?
SSSSSooooollllluuuuutttttiiiiiooooonnnnn:::::
We are given that n = 10,
r = 0.80 and the wrong d 7 should be replaced by 5.
R i
6∑d2
1− i
r R = n ( n2 −1 )
6∑d2
1− i
0.80 = 10 ( 102−1 )
∑d2 = 33
i
Corrected ∑d2 = 33 – 72 + 52 = 9
i
Hence rectified value of rank correlation coefficient
6×9
1−
= 10× ( 102−1 )
= 0.95
(((((ddddd))))) CCCCCOOOOOEEEEEFFFFFFFFFFIIIIICCCCCIIIIIEEEEENNNNNTTTTT OOOOOFFFFF CCCCCOOOOONNNNNCCCCCUUUUURRRRRRRRRREEEEENNNNNTTTTT DDDDDEEEEEVVVVVIIIIIAAAAATTTTTIIIIIOOOOONNNNNSSSSS
A very simple and casual method of finding correlation when we are not serious about the
magnitude of the two variables is the application of concurrent deviations. This method involves
in attaching a positive sign for a x-value (except the first) if this value is more than the previous
value and assigning a negative value if this value is less than the previous value. This is done
for the y-series as well. The deviation in the x-value and the corresponding y-value is known to
be concurrent if both the deviations have the same sign.
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:11)(cid:20)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
Denoting the number of concurrent deviation by c and total number of deviations as m (which
must be one less than the number of pairs of x and y values), the coefficient of concurrent
deviation is given by
( )
−
2c m
±
r = + ............................................................(12.13)
C m
If (2c–m) >0, then we take the positive sign both inside and outside the radical sign and if
(2c–m) <0, we are to consider the negative sign both inside and outside the radical sign.
Like Pearson’s correlation coefficient and Spearman’s rank correlation coefficient, the coefficient
of concurrent deviations also lies between –1 and 1, both inclusive.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111155555 Find the coefficient of concurrent deviations from the following data.
Year : 1990 1991 1992 1993 1994 1995 1996 1997
Price : 25 28 30 23 35 38 39 42
Demand : 35 34 35 30 29 28 26 23
TTTTTaaaaabbbbbllllleeeee 1111122222.....1111100000
SSSSSooooollllluuuuutttttiiiiiooooonnnnn:::::
Computation of Coefficient of Concurrent Deviations.
Year Price Sign of Demand Sign of Product of
deviation deviation from deviation
from the the previous (ab)
previous figure (b)
figure (a)
1990 25 35
1991 28 + 34 – –
1992 30 + 35 + +
1993 23 – 30 – +
1994 35 + 29 – –
1995 38 + 28 – –
1996 39 + 26 – –
1997 42 + 23 – –
In this case, m = number of pairs of deviations = 7
c = No. of positive signs in the product of deviation column = No. of concurrent deviations = 2
(cid:10)(cid:11)(cid:12)(cid:11)(cid:21) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
(2c−m)
Thus r = ± ±
C m
(4−7)
= ± ±
m
(−3)
= ± ±
7
3
= – =−0.65
7
2c−m −3
(Since = we take negative sign both inside and outside of the radical sign)
m 7
Thus there is a negative correlation between price and demand.
1111122222.....55555 RRRRREEEEEGGGGGRRRRREEEEESSSSSSSSSSIIIIIOOOOONNNNN AAAAANNNNNAAAAALLLLLYYYYYSSSSSIIIIISSSSS
In regression analysis, we are concerned with the estimation of one variable for a given value
of another variable (or for a given set of values of a number of variables) on the basis of an
average mathematical relationship between the two variables (or a number of variables).
Regression analysis plays a very important role in the field of every human activity. A
businessman may be keen to know what would be his estimated profit for a given level of
investment on the basis of the past records. Similarly, an outgoing student may like to know
her chance of getting a first class in the final University Examination on the basis of her
performance in the college selection test.
When there are two variables x and y and if y is influenced by x i.e. if y depends on x, then we
get a simple linear regression or simple regression. y is known as dependent variable or regression
or explained variable and x is known as independent variable or predictor or explanator. In
the previous examples since profit depends on investment or performance in the University
Examination is dependent on the performance in the college selection test, profit or performance
in the University Examination is the dependent variable and investment or performance in the
selection test is the In-dependent variable.
In case of a simple regression model if y depends on x, then the regression line of y on x in given
by
y = a + bx …………………… (12.14)
Here a and b are two constants and they are also known as regression parameters. Furthermore,
b is also known as the regression coefficient of y on x and is also denoted by b . We may define
yx
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:11)(cid:22)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
the regression line of y on x as the line of best fit obtained by the method of least squares and
used for estimating the value of the dependent variable y for a known value of the independent
variable x.
The method of least squares involves in minimizing
∑e2 = ∑ (y2 – y)2 = ∑ (y – a – bx)2 ……………………. (12.15)
i i i i i
Where y demotes the actual or observed value and y = a + b , the estimated value of y for a
i i xi i
given value of x, e is the difference between the observed value and the estimated value and e
i i i
is technically known as error or residue. This summation intends over n pairs of observations
of (x y). The line of regression of y or x and the errors of estimation are shown in the following
i, i
figure.
y
Regression line of y on x
0
<
e
n
>0
e
3
y +
bx
2 a
=
y 1 >0 y 2
e
<0 y
e 2
1 y
1
x
0
FIIIIIGGGGGUUUUURRRRREEEEE 1111122222.....77777
SSSSSHHHHHOOOOOWWWWWIIIIINNNNNGGGGG RRRRREEEEEGGGGGRRRRREEEEESSSSSSSSSSIIIIIOOOOONNNNN LLLLLIIIIINNNNNEEEEE OOOOOFFFFF yyyyy OOOOONNNNN xxxxx
AAAAANNNNNDDDDD EEEEERRRRRRRRRROOOOORRRRRSSSSS OOOOOFFFFF EEEEESSSSSTTTTTIIIIIMMMMMAAAAATTTTTIIIIIOOOOONNNNN
Minimisation of (12.15) yields the following equations known as ‘Normal Equations’
. ∑y = na + b∑x ……………….. (12.16)
i i
∑xy = a∑x + b∑ x2 …………..….... (12.17)
i i i i
Solving there two equations for b and a, we have the “least squares” estimates of b and a as
Cov(x,y)
b =
S 2
x
r.S .S
x y
=
S2
x
(cid:10)(cid:11)(cid:12)(cid:11)(cid:23) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
r.S
y
= ..........................................(12.18)
S
x
After estimating b, estimate of a is given by
a=y–bx ......……………………… (12.19)
Substituting the estimates of b and a in (12.14), we get
( ) ( )
y–y r x–x
= ..........................................(12.20)
S S
y x
There may be cases when the variable x depends on y and we may take the regression line of
x on y as
x = a’+ b’y
Unlike the minimization of vertical distances in the scatter diagram as shown in figure (12.7)
for obtaining the estimates of a and b, in this case we minimize the horizontal distances and
get the following normal equation in a’ and b’, the two regression parameters :
∑x = na’ + b’∑y ……………….................. (12.21)
i i
∑xy = a’∑y + b’∑ y2………….............….. (12.22)
i i i i
or solving these equations, we get
cov(x,y) r.S
= x
b’ = b
xy
= S2
y
S
y
..........................(12.23)
anda'=x-b'y …………..................…… (12.24)
A single formula for estimating b is given by
n∑x y −∑x .∑y
i i i i
b = b = ....................(12.25)
yx n∑y2 −(∑y )2
i i
n∑x y −∑x .∑y
i i i i
Similarly, b’ = b = ...........(12.26)
yx n∑y2 −(∑y )2
i i
The standardized form of the regression equation of x on y, as in (12.20), is given by
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:11)(cid:24)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
( )
y–y
x–x
=r …………………................. (12.27)
S S
x y
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111155555 Find the two regression equation from the following data:
x: 2 4 5 5 8 10
y: 6 7 9 10 12 12
Hence estimate y when x is 13 and estimate also x when y is 15.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
TTTTTaaaaabbbbbllllleeeee 1111122222.....1111111111
Computation of Regression Equations
x y x y x 2 y2
i i i i i i
2 6 12 4 36
4 7 28 16 49
5 9 45 25 81
5 10 50 25 100
8 12 96 64 144
10 12 120 100 144
34 56 351 234 554
On the basis of the above table, we have
∑x 34
x= i = =5.6667
n 6
∑y 56
y= i = =9.3333
n 6
∑x y
cov (x, y) = i i −xy
n
351
= −5.6667×9.3333
6
= 58.50–52.8890
= 5.6110
S 2 =
∑x
i
2
−
(
x
)2
x n
(cid:10)(cid:11)(cid:12)(cid:11)(cid:25) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
234
=
−(5.6667)2
6
= 39 – 32.1115
= 6.8885
S 2 =
∑y
i
2
−
(
y
)2
y n
554
=
−(9.3333)2
6
= 92.3333 – 87.1105
= 5.2228
The regression line of y on x is given by
y = a + bx
cov(x,y)
Where b =
S2
x
5.6110
=
6.8885
= 0.8145
and a=y−bx
= 9.3333 – 0.8145 x 5.6667
= 4.7178
Thus the estimated regression equation of y on x is
y = 4.7178 + 0.8145x
When x = 13, the estimated value of y is given by yˆ = 4.7178 + 0.8145 × 13 = 15.3063
The regression line of x on y is given by
x = a’ + b’ y
( )
cov x,y
Where b’ = S 2
y
5.6110
=
5.2228
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:11)(cid:26)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
= 1.0743
and a’ = x–b'y
= 5.6667 – 1.0743 × 9.3333
= – 4.3601
Thus the estimated regression line of x on y is
x = –4.3601 + 1.0743y
When y = 15, the estimate value of x is given by
xˆ = – 4.3601 + 1.0743 × 15
= 11.75
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111166666 Marks of 8 students in Mathematics and statistics are given as:
Mathematics: 80 75 76 69 70 85 72 68
Statistics: 85 65 72 68 67 88 80 70
Find the regression lines. When marks of a student in Mathematics are 90, what are his most
likely marks in statistics?
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
We denote the marks in Mathematics and Statistics by x and y respectively. We are to find the
regression equation of y on x and also of x or y. Lastly, we are to estimate y when x = 90. For
computation advantage, we shift origins of both x and y.
TTTTTaaaaabbbbbllllleeeee 1111122222.....1111122222
Computation of regression lines
Maths Stats u
i
v
i
u
i
v
i
u2 v2
i i
mark (x) mark (y) = x – 74 = y – 76
i i i i
80 85 6 9 54 36 81
75 65 1 –11 –11 1 121
76 72 2 –4 –8 4 16
69 68 –5 –8 40 25 64
70 67 –4 –9 36 16 81
85 88 11 12 132 121 144
72 80 –2 4 –8 4 16
68 70 –6 –6 36 36 36
595 595 3 –13 271 243 559
(cid:10)(cid:11)(cid:12)(cid:20)(cid:27) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
The regression coefficients b (or b ) and b’ (or b ) remain unchanged due to a shift of origin.
yx xy
Applying (12.25) and (12.26), we get
n∑u v −∑u .∑v
i i i i
b = b yx = b vu = n∑u2 −(∑u )2
i i
8.(271)−(3).(−13)
= 8.(243)−(3)2
2168+39
=
1944−9
= 1.1406
n∑u v −∑u .∑v
i i i i
and b’ = b xy = b uv = n∑v2 −(∑v )2
i i
8.(271)−(3).(−13)
=
8.(559)−(−13)2
2168+39
=
4472−169
= 0.5129
Also a = y − bx
(595) (595)
= −1.1406
8 8
= 74.375 – 1.1406 × 74.375
= –10.4571
and a’ = x−b'y
= 74.375– 0.5129 × 74.375
= 36.2280
The regression line of y on x is
y = –10.4571 + 1.1406x
and the regression line of x on y is
x = 36.2281 + 0.5129y
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)(cid:10)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
For x = 90, the most likely value of y is
yˆ = –10.4571 + 1.1406 x 90
= 92.1969
≅ 92
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111177777 The following data relate to the mean and SD of the prices of two shares in a
stock Exchange:
Share Mean (in Rs.) SD (in Rs.)
Company A 44 5.60
Company B 58 6.30
Coefficient of correlation between the share prices = 0.48
Find the most likely price of share A corresponding to a price of Rs. 60 of share B and also the
most likely price of share B for a price of Rs. 50 of share A.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Denoting the share prices of Company A and B respectively by x and y, we are given that
x = Rs. 44 y = Rs. 58
S = Rs. 5.60 S = Rs. 6.30
x y
and r = 0.48
The regression line of y on x is given by
y = a + bx
S
y
r×
Where b =
S
x
6.30
= 0.48×
5.60
= 0.54
a = y − bx
= Rs. (58 – 0.54 × 44)
= Rs. 34.24
Thus the regression line of y on x i.e. the regression line of price of share B on that of share A is
given by
y = Rs. (34.24 + 0.54x)
When x = Rs. 50, = Rs. (34.24 + 0.54 × 50)
(cid:10)(cid:11)(cid:12)(cid:20)(cid:11) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
= Rs. 61.24
= The estimated price of share B for a price of Rs. 50 of share A is Rs.
61.24
Again the regression line of x on y is given by
x = a’ + b’y
S
Where b’ = r× x
S
y
5.60
= 0.48×
6.30
= 0.4267
a = x−b'y
= Rs. (44 – 0.4267 × 58)
= Rs. 19.25
Hence the regression line of x on y i.e. the regression line of price of share A on that of share B
in given by
x = Rs. (19.25 + 0.4267y)
When y = Rs. 60, xˆ = Rs. (19.25 + 0.4267 × 60)
= Rs. 44.85
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111188888 The following data relate the expenditure or advertisement in thousands of
rupees and the corresponding sales in lakhs of rupees.
Expenditure on Ad: 8 10 10 12 15
Sales : 18 20 22 25 28
Find an appropriate regression equation.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
Since sales (y) depend on advertisement (x), the appropriate regression equation is of y on x i.e.
of sales on advertisement. We have, on the basis of the given data,
n = 5, ∑x = 8+10+10+12+15 = 55
∑y = 18+20+22+25+28 = 113
∑xy = 8×18+10×20+10×22+12×25+15×28 = 1284
∑x2 = 82+102+102+122+152 = 633
n∑×y−∑x×∑y
∴ b = n∑x2−(∑x)2
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)(cid:20)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
5×1284−55×113
= 5×633−(55)2
205
=
140
= 1.4643
a = y–bx
113 55
= −1.4643×
5 5
= 22.60 – 16.1073
= 6.4927
Thus, the regression line of y or x i.e. the regression line of sales or advertisement is given by
y = 6.4927 + 1.4643x
1111122222.....66666 PPPPPRRRRROOOOOPPPPPEEEEERRRRRTTTTTIIIIIEEEEESSSSS OOOOOFFFFF RRRRREEEEEGGGGGRRRRREEEEESSSSSSSSSSIIIIIOOOOONNNNN LLLLLIIIIINNNNNEEEEESSSSS
We consider the following important properties of regression lines:
(i) TTTTThhhhheeeee rrrrreeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnntttttsssss rrrrreeeeemmmmmaaaaaiiiiinnnnn uuuuunnnnnccccchhhhhaaaaannnnngggggeeeeeddddd ddddduuuuueeeee tttttooooo aaaaa ssssshhhhhiiiiifffffttttt ooooofffff ooooorrrrriiiiigggggiiiiinnnnn bbbbbuuuuuttttt ccccchhhhhaaaaannnnngggggeeeee ddddduuuuueeeee
tttttooooo aaaaa ssssshhhhhiiiiifffffttttt ooooofffff ssssscccccaaaaallllleeeee.....
This property states that if the original pair of variables is (x, y) and if they are changed to the
pair (u, v) where
x−a y−c
u= andv=
p q
q
×b
b = vu ……………………. (12.28)
yx p
p
× b
and bxy = uv …………………… (12.29)
q
( )
(ii) TTTTThhhhheeeee tttttwwwwwooooo llllliiiiinnnnneeeeesssss ooooofffff rrrrreeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn iiiiinnnnnttttteeeeerrrrrssssseeeeecccccttttt aaaaattttt ttttthhhhheeeee pppppoooooiiiiinnnnnttttt x,y ,,,,, wwwwwhhhhheeeeerrrrreeeee xxxxx aaaaannnnnddddd yyyyy aaaaarrrrreeeee ttttthhhhheeeee vvvvvaaaaarrrrriiiiiaaaaabbbbbllllleeeeesssss
uuuuunnnnndddddeeeeerrrrr cccccooooonnnnnsssssiiiiidddddeeeeerrrrraaaaatttttiiiiiooooonnnnn.....
According to this property, the point of intersection of the regression line of y on x and the
( )
regression line of x on y is x,y i.e. the solution of the simultaneous equations in x and y.
(iii) TTTTThhhhheeeee cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnnttttt ooooofffff cccccooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn bbbbbeeeeetttttwwwwweeeeeeeeeennnnn tttttwwwwwooooo vvvvvaaaaarrrrriiiiiaaaaabbbbbllllleeeeesssss xxxxx aaaaannnnnddddd yyyyy iiiiinnnnn ttttthhhhheeeee sssssiiiiimmmmmpppppllllleeeee gggggeeeeeooooommmmmeeeeetttttrrrrriiiiiccccc
(cid:10)(cid:11)(cid:12)(cid:20)(cid:21) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
mmmmmeeeeeaaaaannnnn ooooofffff ttttthhhhheeeee tttttwwwwwooooo rrrrreeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnntttttsssss..... TTTTThhhhheeeee sssssiiiiigggggnnnnn ooooofffff ttttthhhhheeeee cccccooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnnttttt wwwwwooooouuuuulllllddddd bbbbbeeeee
ttttthhhhheeeee cccccooooommmmmmmmmmooooonnnnn sssssiiiiigggggnnnnn ooooofffff ttttthhhhheeeee tttttwwwwwooooo rrrrreeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn cccccoooooeeeeeffffffffffiiiiiccccciiiiieeeeennnnntttttsssss.....
This property says that if the two regression coefficients are denoted by b (=b) and b (=b’)
yx xy
then the coefficient of correlation is given by
r=± b ×b ………………….. (12.30)
yx xy
If both the regression coefficients are negative, r would be negative and if both are positive, r
would assume a positive value.
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....1111199999 If the relationship between two variables x and u is u + 3x = 10 and between
two other variables y and v is 2y + 5v = 25, and the regression coefficient of y on x is known as
0.80, what would be the regression coefficient of v on u?
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
u + 3x = 10
(x−10/3)
u=
−1/3
and 2y + 5v = 25
(y−25/2)
v=
⇒
−5/2
From (12.28), we have
q
b = ×b
yx vu
p
−5/2
or, 0.80= ×b
vu
−1/3
15
⇒ 0.80= ×b
vu
2
2 8
⇒ b = ×0.80=
vu
15 75
EEEEExxxxxaaaaammmmmpppppllllleeeee 1111122222.....2222200000 For the variables x and y, the regression equations are given as 7x – 3y – 18 = 0
and 4x – y – 11 = 0
(i) Find the arithmetic means of x and y.
(ii) Identify the regression equation of y on x.
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)(cid:22)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
(iii) Compute the correlation coefficient between x and y.
(iv) Given the variance of x is 9, find the SD of y.
SSSSSooooollllluuuuutttttiiiiiooooonnnnn
(i) Since the two lines of regression intersect at the point (x,y), replacing x and y by x and y
respectively in the given regression equations, we get
− −
7x 3y 18=0
and 4x−y−11=0
Solving these two equations, we get x = 3 and y = 1
Thus the arithmetic mean of x and y is given by 3 and 1 respectively.
(ii) Let us assume that 7x – 3y – 18 = 0 represents the regression line of y on x and 4x – y – 11
= 0 represents the regression line of x on y.
Now 7x – 3y – 18 = 0
(7)
⇒ y=(–6)+ x
3
7
∴ b =
yx
3
Again 4x – y – 11 = 0
(11) (1) 1
⇒ x= + y ∴b =
xy
4 4 4
Thus r2 = b × b
yx xy
7 1
= ×
3 4
7
= < 1
12
Since r ≤ 1 ⇒ r2 ≤ 1, our assumptions are correct. Thus, 7x – 3y – 18 = 0 truly represents the
regression line of y on x.
7
(iii) Since r2 =
12
(cid:10)(cid:11)(cid:12)(cid:20)(cid:23) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
7
∴ r = (We take the sign of r as positive since both the regression coefficients are
12
positive)
= 0.7638
S
y
r×
(iv) b =
yx S
x
7 Sy
⇒ = 0.7638× (∴ S 2 = 9 as given)
3 3 x
7
⇒ S =
y 0.7638
= 9.1647
1111122222.....77777 RRRRREEEEEVVVVVIIIIIEEEEEWWWWW OOOOOFFFFF CCCCCOOOOORRRRRRRRRREEEEELLLLLAAAAATTTTTIIIIIOOOOONNNNN AAAAANNNNNDDDDD RRRRREEEEEGGGGGRRRRREEEEESSSSSSSSSSIIIIIOOOOONNNNN AAAAANNNNNAAAAALLLLLYYYYYSSSSSIIIIISSSSS
So far we have discussed the different measures of correlation and also how to fit regression
lines applying the method of ‘Least Squares’. It is obvious that we take recourse to correlation
analysis when we are keen to know whether two variables under study are associated or
correlated and if correlated, what is the strength of correlation. The best measure of correlation
is provided by Pearson’s correlation coefficient. However, one severe limitation of this correlation
coefficient, as we have already discussed, is that it is applicable only in case of a linear
relationship between the two variables.
If two variables x and y are independent or uncorrelated then obviously the correlation
coefficient between x and y is zero. However, the converse of this statement is not necessarily
true i.e. if the correlation coefficient, due to Pearson, between two variables comes out to be
zero, then we cannot conclude that the two variables are independent. All that we can conclude
is that no linear relationship exists between the two variables. This, however, does not rule out
the existence of some non linear relationship between the two variables. For example, if we
consider the following pairs of values on two variables x and y.
(–2, 4), (–1, 1), (0, 0), (1, 1) and (2, 4), then cov (x, y) = (–2+ 4) + (–1+1) + (0×0) + (1×1) + (2×4) = 0
as x = 0
Thus r = 0
xy
This does not mean that x and y are independent. In fact the relationship between x and y is
y = x2. Thus it is always wiser to draw a scatter diagram before reaching conclusion about the
existence of correlation between a pair of variables.
There are some cases when we may find a correlation between two variables although the two
variables are not causally related. This is due to the existence of a third variable which is
related to both the variables under consideration. Such a correlation is known as spurious
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)(cid:24)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
correlation or non-sense correlation. As an example, there could be a positive correlation between
production of rice and that of iron in India for the last twenty years due to the effect of a third
variable time on both these variables. It is necessary to eliminate the influence of the third
variable before computing correlation between the two original variables.
Correlation coefficient measuring a linear relationship between the two variables indicates the
amount of variation of one variable accounted for by the other variable. A better measure for
this purpose is provided by the square of the correlation coefficient, Known as ‘coefficient of
determination’. This can be interpreted as the ratio between the explained variance to total
variance i.e.
Explainedvariance
r2=
Totalvariance
Thus a value of 0.6 for r indicates that (0.6)2 × 100% or 36 per cent of the variation has been
accounted for by the factor under consideration and the remaining 64 per cent variation is due
to other factors. The ‘coefficient of non-determination’ is given by (1–r2) and can be interpreted
as the ratio of unexplained variance to the total variance.
Coefficient of non-determination = (1–r2)
Regression analysis, as we have already seen, is concerned with establishing a functional
relationship between two variables and using this relationship for making future projection.
This can be applied, unlike correlation for any type of relationship linear as well as curvilinear.
TTTTThhhhheeeee tttttwwwwwooooo llllliiiiinnnnneeeeesssss ooooofffff rrrrreeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn cccccoooooiiiiinnnnnccccciiiiidddddeeeee iiiii.....eeeee..... bbbbbeeeeecccccooooommmmmeeeee iiiiidddddeeeeennnnntttttiiiiicccccaaaaalllll wwwwwhhhhheeeeennnnn rrrrr ===== –––––11111 ooooorrrrr 11111 ooooorrrrr iiiiinnnnn ooooottttthhhhheeeeerrrrr
wwwwwooooorrrrrdddddsssss,,,,, ttttthhhhheeeeerrrrreeeee iiiiisssss aaaaa pppppeeeeerrrrrfffffeeeeecccccttttt nnnnneeeeegggggaaaaatttttiiiiivvvvveeeee ooooorrrrr pppppooooosssssiiiiitttttiiiiivvvvveeeee cccccooooorrrrrrrrrreeeeelllllaaaaatttttiiiiiooooonnnnn bbbbbeeeeetttttwwwwweeeeeeeeeennnnn ttttthhhhheeeee tttttwwwwwooooo vvvvvaaaaarrrrriiiiiaaaaabbbbbllllleeeeesssss uuuuunnnnndddddeeeeerrrrr
dddddiiiiissssscccccuuuuussssssssssiiiiiooooonnnnn iiiiifffff rrrrr ===== 00000 RRRRReeeeegggggrrrrreeeeessssssssssiiiiiooooonnnnn llllliiiiinnnnneeeeesssss aaaaarrrrreeeee pppppeeeeerrrrrpppppeeeeennnnndddddiiiiicccccuuuuulllllaaaaarrrrr tttttooooo eeeeeaaaaaccccchhhhh ooooottttthhhhheeeeerrrrr.....
(cid:10)(cid:11)(cid:12)(cid:20)(cid:25) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
EEEEEXXXXXEEEEERRRRRCCCCCIIIIISSSSSEEEEE
SSSSSeeeeettttt AAAAA
Write the correct answers. Each question carries 1 mark.
1. Bivariate Data are the data collected for
(a) Two variables
(b) More than two variables
(c) Two variables at the same point of time
(d) Two variables at different points of time.
2. For a bivariate frequency table having (p + q) classification the total number of cells is
(a) p (b) p + q
(c) q (d) pq
3. Some of the cell frequencies in a bivariate frequency table may be
(a) Negative (b) Zero
(c) a or b (d) Non of these
4. For a p x q bivariate frequency table, the maximum number of marginal distributions is
(a) p (b) p + q
(c) 1 (d) 2
5. For a p x q classification of bivariate data, the maximum number of conditional distributions
is
(a) p (b) p + q
(c) pq (d) p or q
6. Correlation analysis aims at
(a) Predicting one variable for a given value of the other variable
(b) Establishing relation between two variables
(c) Measuring the extent of relation between two variables
(d) Both (b) and (c).
7. Regression analysis is concerned with
(a) Establishing a mathematical relationship between two variables
(b) Measuring the extent of association between two variables
(c) Predicting the value of the dependent variable for a given value of the independent
variable
(d) Both (a) and (c).
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:20)(cid:26)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
8. What is spurious correlation?
(a) It is a bad relation between two variables.
(b) It is very low correlation between two variables.
(c) It is the correlation between two variables having no causal relation.
(d) It is a negative correlation.
9. Scatter diagram is considered for measuring
(a) Linear relationship between two variables
(b) Curvilinear relationship between two variables
(c) Neither (a) nor (b)
(d) Both (a) and (b).
10. If the plotted points in a scatter diagram lie from upper left to lower right, then the
correlation is
(a) Positive (b) Zero
(c) Negative (d) None of these.
11. If the plotted points in a scatter diagram are evenly distributed, then the correlation is
(a) Zero (b) Negative
(c) Positive (d) (a) or (b).
12. If all the plotted points in a scatter diagram lie on a single line, then the correlation is
(a) Perfect positive (b) Perfect negative
(c) Both (a) and (b) (d) Either (a) or (b).
13. The correlation between shoe-size and intelligence is
(a) Zero (b) Positive
(c) Negative (d) None of these.
14. The correlation between the speed of an automobile and the distance travelled by it after
applying the brakes is
(a) Negative (b) Zero
(c) Positive (d) None of these.
15. Scatter diagram helps us to
(a) Find the nature correlation between two variables
(b) Compute the extent of correlation between two variables
(c) Obtain the mathematical relationship between two variables
(d) Both (a) and (c).
(cid:10)(cid:11)(cid:12)(cid:21)(cid:27) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
16. Pearson’s correlation coefficient is used for finding
(a) Correlation for any type of relation
(b) Correlation for linear relation only
(c) Correlation for curvilinear relation only
(d) Both (b) and (c).
17. Product moment correlation coefficient is considered for
(a) Finding the nature of correlation
(b) Finding the amount of correlation
(c) Both (a) and (b)
(d) Either (a) and (b).
18. If the value of correlation coefficient is positive, then the points in a scatter diagram tend
to cluster
(a) From lower left corner to upper right corner
(b) From lower left corner to lower right corner
(c) From lower right corner to upper left corner
(d) From lower right corner to upper right corner.
19. When v = 1, all the points in a scatter diagram would lie
(a) On a straight line directed from lower left to upper right
(b) On a straight line directed from upper left to lower right
(c) On a straight line
(d) Both (a) and (b).
20. Product moment correlation coefficient may be defined as the ratio of
(a) The product of standard deviations of the two variables to the covariance between
them
(b) The covariance between the variables to the product of the variances of them
(c) The covariance between the variables to the product of their standard deviations
(d) Either (b) or (c).
21. The covariance between two variables is
(a) Strictly positive (b) Strictly negative
(c) Always 0 (d) Either positive or negative or zero.
22. The coefficient of correlation between two variables
(a) Can have any unit.
(b) Is expressed as the product of units of the two variables
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:21)(cid:10)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
(c) Is a unit free measure
(d) None of these.
23. What are the limits of the correlation coefficient?
(a) No limit (b) –1 and 1
(c) 0 and 1, including the limits (d) –1 and 1, including the limits
24. In case the correlation coefficient between two variables is 1, the relationship between the
two variables would be
(a) y = a + bx (b) y = a + bx, b > 0
(c) y = a + bx, b < 0 (d) y = a + bx, both a and b being positive.
25. If the relationship between two variables x and y in given by 2x + 3y + 4 = 0, then the
value of the correlation coefficient between x and y is
(a) 0 (b) 1
(c) –1 (d) negative.
26. For finding correlation between two attributes, we consider
(a) Pearson’s correlation coefficient
(b) Scatter diagram
(c) Spearman’s rank correlation coefficient
(d) Coefficient of concurrent deviations.
27. For finding the degree of agreement about beauty between two Judges in a Beauty Contest,
we use
(a) Scatter diagram (b) Coefficient of rank correlation
(c) Coefficient of correlation (d) Coefficient of concurrent deviation.
28. If there is a perfect disagreement between the marks in Geography and Statistics, then
what would be the value of rank correlation coefficient?
(a) Any value (b) Only 1
(c) Only –1 (d) (b) or (c)
29. When we are not concerned with the magnitude of the two variables under discussion,
we consider
(a) Rank correlation coefficient (b) Product moment correlation coefficient
(c) Coefficient of concurrent deviation (d) (a) or (b) but not (c).
30. What is the quickest method to find correlation between two variables?
(a) Scatter diagram (b) Method of concurrent deviation
(c) Method of rank correlation (d) Method of product moment correlation
(cid:10)(cid:11)(cid:12)(cid:21)(cid:11) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
31. What are the limits of the coefficient of concurrent deviations?
(a) No limit
(b) Between –1 and 0, including the limiting values
(c) Between 0 and 1, including the limiting values
(d) Between –1 and 1, the limiting values inclusive
32. If there are two variables x and y, then the number of regression equations could be
(a) 1 (b) 2
(c) Any number (d) 3.
33. Since Blood Pressure of a person depends on age, we need consider
(a) The regression equation of Blood Pressure on age
(b) The regression equation of age on Blood Pressure
(c) Both (a) and (b)
(d) Either (a) or (b).
34. The method applied for deriving the regression equations is known as
(a) Least squares (b) Concurrent deviation
(c) Product moment (d) Normal equation.
35. The difference between the observed value and the estimated value in regression analysis
is known as
(a) Error (b) Residue
(c) Deviation (d) (a) or (b).
36. The errors in case of regression equations are
(a) Positive (b) Negative
(c) Zero (d) All these.
37. The regression line of y on is derived by
(a) The minimisation of vertical distances in the scatter diagram
(b) The minimisation of horizontal distances in the scatter diagram
(c) Both (a) and (b)
(d) (a) or (b).
38. The two lines of regression become identical when
(a) r = 1 (b) r = –1
(c) r = 0 (d) (a) or (b).
39. What are the limits of the two regression coefficients?
(a) No limit (b) Must be positive
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:21)(cid:20)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
(c) One positive and the other negative
(d) Product of the regression coefficient must be numerically less than unity.
40. The regression coefficients remain unchanged due to a
(a) Shift of origin (b) Shift of scale
(c) Both (a) and (b) (d) (a) or (b).
41. If the coefficient of correlation between two variables is –0 9, then the coefficient of
determination is
(a) 0.9 (b) 0.81
(c) 0.1 (d) 0.19.
42. If the coefficient of correlation between two variables is 0.7 then the percentage of variation
unaccounted for is
(a) 70% (b) 30%
(c) 51% (d) 49%
SSSSSeeeeettttt BBBBB
Answer the following questions by writing the correct answers. Each question carries 2 marks.
1. If for two variable x and y, the covariance, variance of x and variance of y are 40, 16 and
256 respectively, what is the value of the correlation coefficient?
(a) 0.01 (b) 0.625
(c) 0.4 (d) 0.5
2. If cov(x, y) = 15, what restrictions should be put for the standard deviations of x and y?
(a) No restriction.
(b) The product of the standard deviations should be more than 15.
(c) The product of the standard deviations should be less than 15.
(d) The sum of the standard deviations should be less than 15.
3. If the covariance between two variables is 20 and the variance of one of the variables is 16,
what would be the variance of the other variable?
(a) More than 100 (b) More than 10
(c) Less than 10 (d) More than 1.25
4. If y = a + bx, then what is the coefficient of correlation between x and y?
(a) 1 (b) –1
(c) 1 or –1 according as b > 0 or b < 0 (d) none of these.
5. If g = 0.6 then the coefficient of non-determination is
(a) 0.4 (b) –0.6
(c) 0.36 (d) 0.64
(cid:10)(cid:11)(cid:12)(cid:21)(cid:21) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
6. If u + 5x = 6 and 3y – 7v = 20 and the correlation coefficient between x and y is 0.58 then
what would be the correlation coefficient between u and v?
(a) 0.58 (b) –0.58
(c) –0.84 (d) 0.84
7. If the relation between x and u is 3x + 4u + 7 = 0 and the correlation coefficient between x
and y is –0.6, then what is the correlation coefficient between u and y?
(a) –0.6 (b) 0.8
(c) 0.6 (d) –0.8
8 From the following data
x: 2 3 5 4 7
y: 4 6 7 8 10
Two coefficient of correlation was found to be 0.93. What is the correlation between u
and v as given below?
u: –3 –2 0 –1 2
v: –4 –2 –1 0 2
(a) –0.93 (b) 0.93 (c) 0.57 (d) –0.57
9. Referring to the data presented in Q. No. 8, what would be the correlation between u and
v?
u: 10 15 25 20 35
v: –24 –36 –42 –48 –60
(a) –0.6 (b) 0.6 (c) –0.93 (d) 0.93
10. If the sum of squares of difference of ranks, given by two judges A and B, of 8 students in
21, what is the value of rank correlation coefficient?
(a) 0.7 (b) 0.65 (c) 0.75 (d) 0.8
11. If the rank correlation coefficient between marks in management and mathematics for a
group of student in 0.6 and the sum of squares of the differences in ranks in 66, what is
the number of students in the group?
(a) 10 (b) 9 (c) 8 (d) 11
12. While computing rank correlation coefficient between profit and investment for the last 6
years of a company the difference in rank for a year was taken 3 instead of 4. What is the
rectified rank correlation coefficient if it is known that the original value of rank correlation
coefficient was 0.4?
(a) 0.3 (b) 0.2 (c) 0.25 (d) 0.28
13. For 10 pairs of observations, No. of concurrent deviations was found to be 4. What is the
value of the coefficient of concurrent deviation?
(a) 0.2 (b) – 0.2 (c) 1/3 (d) –1/3
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:21)(cid:22)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
14. The coefficient of concurrent deviation for p pairs of observations was found to be 1/ 3
. If the number of concurrent deviations was found to be 6, then the value of p is.
(a) 10 (b) 9 (c) 8 (d) none of these
15. What is the value of correlation coefficient due to Pearson on the basis of the following
data:
x: –5 –4 –3 –2 –1 0 1 2 3 4 5
y: 27 18 11 6 3 2 3 6 11 18 27
(a) 1 (b) –1 (c) 0 (d) –0.5
16. Following are the two normal equations obtained for deriving the regression line of
y and x:
5a + 10b = 40
10a + 25b = 95
The regression line of y on x is given by
(a) 2x + 3y = 5 (b) 2y + 3x = 5 (c) y = 2 + 3x (d) y = 3 + 5x
17. If the regression line of y on x and of x on y are given by 2x + 3y = –1 and 5x + 6y = –1 then
the arithmetic means of x and y are given by
(a) (1, –1) (b) (–1, 1) (c) (–1, –1) (d) (2, 3)
18. Given the regression equations as 3x + y = 13 and 2x + 5y = 20, which one is the regression
equation of y on x?
(a) 1st equation (b) 2nd equation (c) both (a) and (b) (d) none of these.
19. Given the following equations: 2x – 3y = 10 and 3x + 4y = 15, which one is the regression
equation of x on y ?
(a) 1st equation (b) 2nd equation (c) both the equations (d) none of these
20. If u = 2x + 5 and v = –3y – 6 and regression coefficient of y on x is 2.4, what is the
regression coefficient of v on u?
(a) 3.6 (b) –3.6 (c) 2.4 (d) –2.4
21. If 4y – 5x = 15 is the regression line of y on x and the coefficient of correlation between x
and y is 0.75, what is the value of the regression coefficient of x on y?
(a) 0.45 (b) 0.9375 (c) 0.6 (d) none of these
22. If the regression line of y on x and that of x on y are given by y = –2x + 3 and 8x = –y + 3
respectively, what is the coefficient of correlation between x and y?
(a) 0.5 (b) –1/ 2 (c) –0.5 (d) none of these
23. If the regression coefficient of y on x, the coefficient of correlation between x and y and
variance of y are –3/4, – 3/2 and 4 respectively, what is the variance of x?
(a) 2/ 3/2 (b) 16/3 (c) 4/3 (d) 4
(cid:10)(cid:11)(cid:12)(cid:21)(cid:23) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
24. If y = 3x + 4 is the regression line of y on x and the arithmetic mean of x is –1, what is the
arithmetic mean of y?
(a) 1 (b) –1 (c) 7 (d) none of these
SSSSSEEEEETTTTT CCCCC
Write down the correct answers. Each question carries 5 marks.
1. What is the coefficient of correlation from the following data?
x: 1 2 3 4 5
y: 8 6 7 5 5
(a) 0.75 (b) –0.75 (c) –0.85 (d) 0.82
2. The coefficient of correlation between x and y where
x: 64 60 67 59 69
y: 57 60 73 62 68
is
(a) 0.655 (b) 0.68 (c) 0.73 (d) 0.758
3. What is the coefficient of correlation between the ages of husbands and wives from the
following data?
Age of husband (year): 46 45 42 40 38 35 32 30 27 25
Age of wife (year): 37 35 31 28 30 25 23 19 19 18
(a) 0.58 (b) 0.98 (c) 0.89 (d) 0.92
4. Given that for twenty pairs of observations, ∑xu = 525, ∑x = 129, ∑u = 97, ∑x2 = 687,
∑u2 = 427 and y = 10 – 3u, the coefficient of correlation between x and y is
(a) –0.7 (b) 0.74 (c) –0.74 (d) 0.75
5. The following results relate to bivariate date on (x, y):
∑xy=
414,
∑x=
120,
∑y
= 90,
∑x2
= 600,
∑y2
= 300, n = 30, later or, it was known that
two pairs of observations (12, 11) and (6, 8) were wrongly taken, the correct pairs of
observations being (10, 9) and (8, 10). The corrected value of the correlation coefficient is
(a) 0.752 (b) 0.768 (c) 0.846 (d) 0.953
6. The following table provides the distribution of items according to size groups and also
the number of defectives:
Size group: 9-11 11-13 13-15 15-17 17-19
No. of items: 250 350 400 300 150
No. of defective items: 25 70 60 45 20
The correlation coefficient between size and defectives is
(a) 0.25 (b) 0.12 (c) 0.14 (d) 0.07
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:21)(cid:24)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
7. For two variables x and y, it is known that cov (x, y) = 80, variance of x is 16 and sum of
squares of deviation of y from its mean is 250. The number of observations for this bivariate
data is
(a) 7 (b) 8 (c) 9 (d) 10
8. Eight contestants in a musical contest were ranked by two judges A and B in the following
manner:
Serial Number
of the contestants: 1 2 3 4 5 6 7 8
Rank by Judge A: 7 6 2 4 5 3 1 8
Rank by Judge B: 5 4 6 3 8 2 1 7
The rank correlation coefficient is
(a) 0.65 (b) 0.63 (c) 0.60 (d) 0.57
9. Following are the marks of 10 students in Botany and Zoology:
Serial No.: 1 2 3 4 5 6 7 8 9 10
Marks in
Botany: 58 43 50 19 28 24 77 34 29 75
Marks in
Zoology: 62 63 79 56 65 54 70 59 55 69
The coefficient of rank correlation between marks in Botany and Zoology is
(a) 0.65 (b) 0.70 (c) 0.72 (d) 0.75
10. What is the value of Rank correlation coefficient between the following marks in Physics
and Chemistry:
Roll No.: 1 2 3 4 5 6
Marks in Physics: 25 30 46 30 55 80
Marks in Chemistry: 30 25 50 40 50 78
(a) 0.782 (b) 0.696 (c) 0.932 (d) 0.857
11. What is the coefficient of concurrent deviations for the following data:
Supply: 68 43 38 78 66 83 38 23 83 63 53
Demand: 65 60 55 61 35 75 45 40 85 80 85
(a) 0.82 (b) 0.85 (c) 0.89 (d) –0.81
12. What is the coefficient of concurrent deviations for the following data:
Year: 1996 1997 1998 1999 2000 2001 2002 2003
Price: 35 38 40 33 45 48 49 52
Demand: 36 35 31 36 30 29 27 24
(a) –0.43 (b) 0.43 (c) 0.5 (d) 2
(cid:10)(cid:11)(cid:12)(cid:21)(cid:25) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
13. The regression equation of y on x for the following data:
x 41 82 62 37 58 96 127 74 123 100
y 28 56 35 17 42 85 105 61 98 73
Is given by
(a) y = 1.2x – 15 (b) y = 1.2x + 15 (c) y = 0.93x – 14.64 (d) y = 1.5x – 10.89
14. The following data relate to the heights of 10 pairs of fathers and sons:
(175, 173), (172, 172), (167, 171), (168, 171), (172, 173), (171, 170), (174, 173), (176, 175) (169, 170), (170, 173)
The regression equation of height of son on that of father is given by
(a) y = 100 + 5x (b) y = 99.708 + 0.405x (c) y = 89.653 + 0.582x (d) y = 88.758 + 0.562x
15. The two regression coefficients for the following data:
x: 38 23 43 33 28
y: 28 23 43 38 8
are
(a) 1.2 and 0.4 (b) 1.6 and 0.8 (c) 1.7 and 0.8 (d) 1.8 and 0.3
16. For y = 25, what is the estimated value of x, from the following data:
X: 11 12 15 16 18 19 21
Y: 21 15 13 12 11 10 9
(a) 15 (b) 13.926 (c) 13.588 (d) 14.986
17. Given the following data:
Variable: x y
Mean: 80 98
Variance: 4 9
Coefficient of correlation = 0.6
What is the most likely value of y when x = 90 ?
(a) 90 (b) 103 (c) 104 (d) 107
18. The two lines of regression are given by
8x + 10y = 25 and 16x + 5y = 12 respectively.
If the variance of x is 25, what is the standard deviation of y?
(a) 16 (b) 8 (c) 64 (d) 4
19. Given below the information about the capital employed and profit earned by a company
over the last twenty five years:
Mean SD
Capital employed ( 0000 Rs) 62 5
Profit earned ( 000 Rs) 25 6
(cid:19)(cid:5)(cid:3)(cid:5)(cid:17)(cid:19)(cid:5)(cid:17)(cid:1)(cid:19) (cid:10)(cid:11)(cid:12)(cid:21)(cid:26)
Copyright -The Institute of Chartered Accountants of India
(cid:1)(cid:2)(cid:3)(cid:3)(cid:4)(cid:5)(cid:6)(cid:7)(cid:8)(cid:2)(cid:9)(cid:10)(cid:6)(cid:9)(cid:11)(cid:10)(cid:3)(cid:4)(cid:12)(cid:3)(cid:4)(cid:13)(cid:13)(cid:8)(cid:2)(cid:9)
Correlation Coefficient between capital and profit = 0.92. The sum of the Regression
coefficients for the above data would be:
(a) 1.871 (b) 2.358 (c) 1.968 (d) 2.346
20. The coefficient of correlation between cost of advertisement and sales of a product on the
basis of the following data:
Ad cost (000 Rs): 75 81 85 105 93 113 121 125
Sales (000 000 Rs): 35 45 59 75 43 79 87 95
is
(a) 0.85 (b) 0.89 (c) 0.95 (d) 0.98
AAAAANNNNNSSSSSWWWWWEEEEERRRRRSSSSS
SSSSSeeeeettttt AAAAA
1. (c) 2. (d) 3. (b) 4. (d) 5. (b) 6. (d)
7. (d) 8. (c) 9. (d) 10. (c) 11. (a) 12. (d)
13. (a) 14. (a) 15. (a) 16. (b) 17. (c) 18. (a)
19. (a) 20. (c) 21. (d) 22. (c) 23. (c) 24. (b)
25. (c) 26. (c) 27. (b) 28. (c) 29. (c) 30. (b)
31. (d) 32. (b) 33. (a) 34. (a) 35. (d) 36. (d)
37. (a) 38. (d) 39. (d) 40. (a) 41. (b) 42. (c)
SSSSSeeeeettttt BBBBB
1. (b) 2. (b) 3. (a) 4. (c) 5. (d) 6. (b)
7. (c) 8. (b) 9. (c) 10. (c) 11. (a) 12. (b)
13. (d) 14. (a) 15. (c) 16. (c) 17. (c) 18. (b)
19. (d) 20. (b) 21. (a) 22. (c) 23. (b) 24. (a)
SSSSSeeeeettttt CCCCC
1. (c) 2. (a) 3. (b) 4. (c) 5. (c) 6. (d)
7. (d) 8. (d) 9. (d) 10. (d) 11. (c) 12. (a)
13. (c) 14. (b) 15. (a) 16. (c) 17. (d) 18. (b)
19. (a) 20. (c)
(cid:10)(cid:11)(cid:12)(cid:22)(cid:27) (cid:1)(cid:13)(cid:14)(cid:14)(cid:13)(cid:15)(cid:8) (cid:4)(cid:7)(cid:13)(cid:16)(cid:17)(cid:1)(cid:17)(cid:6)(cid:15)(cid:1)(cid:18)(cid:8) (cid:5)(cid:6)(cid:19)(cid:5)
Copyright -The Institute of Chartered Accountants of India
Showing pages 1–50 of 58Next 50 pages →