Follow Us
Select Medium / माध्यम चुनें:
Eng (English) Beng (বাংলা) Hindi (हिन्दी)
WBB • Class XI • Economics • Ch 10
Estimated Time: 45 Mins
Study Progress: In Progress

Correlation

Correlation represents one of the most foundational analytical concepts in bivariate statistical economics, providing a rigorous quantitative methodology to investigate, measure, and interpret the degree and direction of mutual association between two or more interrelated socioeconomic variables. Prescribed under the West Bengal Council of Higher Secondary Education (WBCHSE) Class 11 Economics curriculum within the Quantitative Economics and Statistics syllabus, this chapter bridges descriptive univariate summaries—such as measures of central tendency and dispersion—with multivariate econometric modeling. While univariate statistics evaluate the distribution of a single isolated variable, real-world economic phenomena operate in deeply interconnected systems where variables continuously co-vary: market price fluctuates alongside consumer demand, agricultural yield varies with seasonal rainfall, household consumption expenditure expands with disposable income, and corporate investment shifts with prevailing interest rates. The study of correlation equips students with the analytical apparatus to determine whether two economic series move synchronously in the same direction (positive correlation), diverge in opposite directions (negative correlation), or exhibit complete linear independence (zero correlation). The curriculum systematically examines qualitative graphical techniques through scatter diagrams, quantitative parametric measurement via Karl Pearson's product-moment correlation coefficient (r), and non-parametric ordinal analysis through Charles Spearman's rank correlation coefficient (R). Crucially, the chapter emphasizes the vital epistemological distinction between statistical correlation and causal determination, warning students against spurious or nonsense associations where apparent mathematical co-movements arise from common external trends rather than substantive causal mechanisms.

Why This Chapter Matters

In contemporary empirical economics, financial market analysis, and public policy formulation, correlation analysis serves as the primary exploratory diagnostic for hypothesis formulation and econometric forecasting. Government planners and central bankers scrutinize correlation coefficients to assess the empirical transmission mechanisms of fiscal and monetary interventions. For instance, the Reserve Bank of India analyzes the correlation between benchmark policy interest rates and commercial bank lending volumes to evaluate monetary transmission, while macroeconomists investigate the historical relationship between fiscal deficits and consumer price inflation. In financial economics, modern portfolio theory, pioneered by Nobel laureate Harry Markowitz, rests entirely upon the correlation matrix of asset returns: combining stocks and bonds with negative or low correlation enables institutional and retail investors to minimize overall portfolio risk without sacrificing expected returns. Similarly, business enterprises rely on correlation metrics to estimate price elasticities of demand, calculate cross-price effects of rival products, and evaluate the return on marketing expenditures. For higher secondary economics students, mastering the algebraic properties, computational pathways, and interpretative boundaries of Karl Pearson's and Spearman's correlation coefficients establishes the indispensable quantitative foundation required for advanced regression analysis, econometrics, and data-driven economic research.

Chapter Roadmap & Progression

1 Conceptual Foundations, Meaning & C...
2 Graphical Analysis: The Scatter Dia...
3 Karl Pearson's Coefficient of Linea...
4 Fundamental Mathematical Properties...
5 Spearman's Rank Correlation Coeffic...
6 Empirical Economic Applications, Li...

Complete Concept Guide (100% Curriculum Coverage)

Conceptual Foundations, Meaning & Classification of Correlation

1. Meaning & Definition of Correlation

In univariate statistics, analysis is confined to the characteristics of a single variable (such as national income, factory wages, or wheat prices). However, in economic life, variables rarely exist in isolation; they continuously interact. When two variables vary together in such a manner that a change in one is accompanied by an equivalent or systematic change in the other, they are said to be correlated.

Prominent statisticians have defined correlation as follows:

  • Croxton & Cowden: "When the relationship is of a quantitative nature, the appropriate statistical tool for discovering and measuring the relationship and expressing it in a brief formula is known as correlation."
  • A. M. Tuttle: "Correlation is an analysis of the covariation between two or more variables."
  • Boddington: "Whenever some definite connection exists between two or more groups, classes, or series of data, there is said to be correlation."
2. Bivariate Distributions & Joint Variation

Data involving two simultaneous characteristics measured on the same statistical unit (e.g., height and weight of an individual, price and quantity demanded of a commodity, fertilizer dosage and crop yield per acre) constitute a bivariate distribution, represented as pairs of observations $(x_1, y_1), (x_2, y_2), \dots, (x_n, y_n)$. Correlation measures the strength and direction of linear association within such paired observations.

3. Comprehensive Classification of Correlation

Correlation is classified according to three distinct analytical dimensions:

Classification Dimension Category Type Nature of Association Concrete Economic Examples
I. Direction of Movement Positive (Direct) Correlation Both variables move in the same direction: an increase in $X$ is accompanied by an increase in $Y$, or a decrease in $X$ leads to a decrease in $Y$. Price and Quantity Supplied ($P \uparrow, Q_S \uparrow$); Household Income and Consumption Expenditure ($Y \uparrow, C \uparrow$); Advertising Spend and Sales Revenue.
Negative (Inverse) Correlation Variables move in opposite directions: an increase in $X$ is accompanied by a decrease in $Y$, and vice versa. Price and Quantity Demanded ($P \uparrow, Q_D \downarrow$); Winter Temperature and Woollen Garment Sales; Speed of Production and Delivery Time.
II. Ratio / Constancy of Change Linear Correlation The amount of change in one variable bears a constant ratio to the amount of change in the other variable. When plotted on a graph, all observations form a straight line ($Y = a + bX$). Cost of raw material where unit price is strictly fixed (e.g., total cost increases by exactly ₹50 for every additional kg purchased).
Non-Linear (Curvilinear) Correlation The ratio of change between variables is not constant but varies across different ranges of data. When plotted, the points trace a curve (parabolic, hyperbolic, logarithmic). Law of Variable Proportions (Output initially increases at an increasing rate, then at a diminishing rate, and eventually declines); Engel Curves for luxury goods.
III. Number of Variables Simple Correlation Study of the mutual relationship between exactly two variables ($X$ and $Y$). Relationship between money supply and the general wholesale price level.
Multiple Correlation Simultaneous study of the relationship among three or more variables together. Study of total agricultural output of paddy ($Y$) as determined jointly by rainfall ($X_1$), fertilizer application ($X_2$), and seed quality ($X_3$).
Partial Correlation Study of the relationship between two specific variables while holding the effects of other related variables constant. Examining the relationship between rainfall ($X_1$) and crop yield ($Y$) while keeping fertilizer usage ($X_2$) constant.
4. Epistemological Distinction: Correlation vs. Causation
Crucial Board Principle: Correlation Does NOT Imply Causation!
Correlation simply establishes numerical co-variation; it does not prove that one variable is the cause and the other is the effect. A high correlation coefficient may exist under four entirely distinct circumstances:
  1. Direct Causation: $X$ causes $Y$ (e.g., excessive money supply causes demand-pull inflation).
  2. Mutual / Reciprocal Causation: $X$ affects $Y$, and $Y$ simultaneously affects $X$ (e.g., price and demand in general equilibrium, investment and income via the multiplier-accelerator interaction).
  3. Common Underlying Cause (Lurking Variable): Both $X$ and $Y$ are driven by a third external factor $Z$. For instance, agricultural wage rates and school attendance may both rise during periods of bumper harvests due to overall rural prosperity.
  4. Spurious or Nonsense Correlation: High mathematical correlation arising purely by sheer historical coincidence or secular trends with zero logical connection (e.g., correlation between teacher salaries in India and liquor consumption in England, or stork populations and human birth rates in Europe). Such correlations are economically meaningless.

Graphical Analysis: The Scatter Diagram Method

1. Concept & Technique of Scatter Diagrams

The Scatter Diagram (or Scattergram) is the simplest, most intuitive graphical technique for visually inspecting the nature, direction, and strength of correlation between two variables. It requires no complex mathematical computations and provides an immediate visual summary of bivariate relationships.

Construction Procedure:

  1. Take the independent or base variable on the horizontal axis ($X$-axis) and the dependent or related variable on the vertical axis ($Y$-axis).
  2. Select convenient and proportionate scales for both axes so that the entire range of data is clearly distributed across the graph.
  3. Plot each pair of values $(x_i, y_i)$ as a discrete coordinate dot $(\bullet)$ on the Cartesian plane.
  4. The resultant cluster of plotted dots constitutes the scatter diagram. The general shape, slope, and dispersion of the dot cloud reveal the underlying correlation pattern.
2. Visual Interpretation of Scatter Patterns
Diagrammatic Pattern Visual Appearance of Dots Degree & Direction Correlation Value
Perfect Positive Correlation All dots lie precisely on a single straight line rising upward from the lower-left corner to the upper-right corner. Perfect Direct $$r = +1$$
High Degree of Positive Correlation Dots do not form a strict line but cluster tightly along an upward-sloping narrow band from southwest to northeast. Strong Direct $$+0.75 \le r < +1$$
Low Degree of Positive Correlation Dots show a discernible upward tilt but are widely scattered and loosely dispersed. Weak Direct $$0 < r < +0.25$$
Perfect Negative Correlation All dots lie precisely on a single straight line sloping downward from the upper-left corner to the lower-right corner. Perfect Inverse $$r = -1$$
High Degree of Negative Correlation Dots cluster tightly along a downward-sloping band from northwest to southeast. Strong Inverse $$-1 < r \le -0.75$$
Zero Correlation (Absence of Linear Association) Dots are scattered completely haphazardly across the diagram with no discernible upward or downward trend, forming a circular cloud or horizontal/vertical band. No Linear Association $$r = 0$$
Curvilinear Relationship Dots trace a well-defined curved path (such as an inverted U-shape or parabola). While a strong non-linear functional relation exists, the linear correlation coefficient is zero. Non-Linear Association $$r \approx 0 \text{ (linear)}$$
3. Critical Evaluation: Merits & Limitations
  • Merits:
    • Extremely simple and non-mathematical; easily understood by non-statisticians.
    • Instantly reveals whether the association is linear or non-linear, and whether positive or negative.
    • Visual inspection immediately exposes extreme outliers or recording anomalies that could distort algebraic averages.
  • Limitations:
    • Purely qualitative and descriptive; it does not provide an exact numerical coefficient indicating the precise magnitude of correlation.
    • Subjective visual interpretation: different observers may judge the tightness of the scatter band differently.
    • Incapable of measuring multivariate relationships involving more than two variables simultaneously.

Karl Pearson's Coefficient of Linear Correlation (Product-Moment r)

1. Theoretical Foundation & Mathematical Definition

Formulated by the eminent British biometrician and statistician Karl Pearson in 1896, the Product-Moment Correlation Coefficient (denoted by the symbol $r$ or $r_{xy}$) is the most widely utilized mathematical measure of linear association between two quantitative variables.

Pearson defined $r$ as the ratio of the covariance between variables $X$ and $Y$ to the product of their individual standard deviations:

$$r = \frac{\text{Cov}(X, Y)}{\sigma_X \cdot \sigma_Y}$$

Where:

  • $\text{Cov}(X, Y) = \frac{1}{N}\sum (X - \bar{X})(Y - \bar{Y}) = \frac{\sum xy}{N}$ (with $x = X - \bar{X}$ and $y = Y - \bar{Y}$)
  • $\sigma_X = \sqrt{\frac{1}{N}\sum (X - \bar{X})^2} = \sqrt{\frac{\sum x^2}{N}}$
  • $\sigma_Y = \sqrt{\frac{1}{N}\sum (Y - \bar{Y})^2} = \sqrt{\frac{\sum y^2}{N}}$

Substituting these into the ratio cancels out the sample size $N$, yielding the classic product-moment formula:

$$r = \frac{\sum (X - \bar{X})(Y - \bar{Y})}{\sqrt{\sum (X - \bar{X})^2} \cdot \sqrt{\sum (Y - \bar{Y})^2}} = \frac{\sum xy}{\sqrt{\sum x^2 \cdot \sum y^2}}$$
2. Computational Pathways for Board Examinations

In WBCHSE examinations, candidates can compute $r$ using three standard methods depending on whether the arithmetic means are whole integers or fractions:

Method Applicability Condition Algebraic Formula
Direct Method (Actual Means) Recommended when both $\bar{X}$ and $\bar{Y}$ are exact whole integers (non-fractional). $$r = \frac{\sum xy}{\sqrt{\sum x^2 \cdot \sum y^2}}$$
where $x = X - \bar{X},\, y = Y - \bar{Y}$
Short-cut Method (Assumed Means) Recommended when actual means $\bar{X}$ and/or $\bar{Y}$ are fractions, avoiding tedious decimal calculations. $$r = \frac{N\sum d_x d_y - (\sum d_x)(\sum d_y)}{\sqrt{N\sum d_x^2 - (\sum d_x)^2} \cdot \sqrt{N\sum d_y^2 - (\sum d_y)^2}}$$
where $d_x = X - A_x,\, d_y = Y - A_y$ ($A_x, A_y$ are arbitrary assumed means).
Step-Deviation Method (Change of Scale) Recommended when data values have common factors/multipliers ($c_x, c_y$), drastically reducing numbers. $$r = \frac{N\sum u v - (\sum u)(\sum v)}{\sqrt{N\sum u^2 - (\sum u)^2} \cdot \sqrt{N\sum v^2 - (\sum v)^2}}$$
where $u = \frac{X - A_x}{c_x},\, v = \frac{Y - A_y}{c_y}$ ($c_x, c_y$ are common class factors).
Raw Data Formula (Without Deviations) Useful for rapid small-sample calculations or pre-aggregated summary statistics. $$r = \frac{N\sum XY - (\sum X)(\sum Y)}{\sqrt{N\sum X^2 - (\sum X)^2} \cdot \sqrt{N\sum Y^2 - (\sum Y)^2}}$$
3. Necessary Computational Checks
Golden Rules for Tabulation:
  • When using the direct method, verify that $\sum x = \sum (X - \bar{X}) = 0$ and $\sum y = \sum (Y - \bar{Y}) = 0$. If these sums do not equal zero, your arithmetic mean calculation is incorrect!
  • In short-cut methods, $\sum d_x$ and $\sum d_y$ will generally be non-zero; always double-check $\sum d_x^2$ and $(\sum d_x)^2$, as confusing these two terms is the most frequent student error.
  • The denominator is the product of two positive square roots; it can never be negative or zero (unless all observations in a variable are identical, in which case variance is zero and $r$ is undefined).

Fundamental Mathematical Properties & Interpretation of Pearson's r

1. Five Cardinal Mathematical Properties of Pearson's r

Karl Pearson's coefficient of correlation satisfies five essential mathematical theorems that frequently form the subject of theoretical questions in Higher Secondary examinations:

Property I: Strict Boundedness between -1 and +1

The numerical value of the correlation coefficient is strictly bounded within the closed interval $[-1, +1]$:

$$-1 \le r \le +1 \quad \text{or} \quad |r| \le 1$$

Mathematical Rationale: This property derives directly from the Cauchy-Schwarz Inequality: for any two real vectors $\vec{u}$ and $\vec{v}$, $(\vec{u} \cdot \vec{v})^2 \le |\vec{u}|^2 |\vec{v}|^2$. Applying this to centered deviations, $[\sum (x-\bar{x})(y-\bar{y})]^2 \le [\sum (x-\bar{x})^2][\sum (y-\bar{y})^2]$, which guarantees that $r^2 \le 1$, meaning $-1 \le r \le +1$. If a calculation yields $r = 1.25$ or $r = -1.40$, an arithmetic error is certain.

Property II: Invariance to Change of Origin and Change of Scale

The value of $r$ is completely independent of the choice of origin (addition or subtraction of constants) and choice of scale (multiplication or division by positive constants).

If original variables $X$ and $Y$ are transformed into new variables $U$ and $V$ by:

$$U = \frac{X - a}{h} \qquad \text{and} \qquad V = \frac{Y - b}{k}$$

where $a, b$ are arbitrary origin shifts and $h, k$ are non-zero scale constants, then:

$$r_{UV} = \left( \frac{h \cdot k}{|h \cdot k|} \right) r_{XY}$$
  • If $h$ and $k$ have the same sign (both positive or both negative), then $r_{UV} = r_{XY}$.
  • If $h$ and $k$ have opposite signs (one positive and one negative), then $r_{UV} = -r_{XY}$.
Property III: Pure Dimensionless Number

Because $r$ is the ratio of covariance (measured in units of $X \times$ units of $Y$) to the product of standard deviations (measured in units of $X \times$ units of $Y$), all physical measurement units cancel out entirely. Thus, $r$ is a pure dimensionless number, permitting direct comparison between relationships measured in completely different dimensions (e.g., height in inches vs. weight in kg, or price in rupees vs. quantity in quintals).

Property IV: Symmetry with Respect to Variables

Correlation is a symmetric relationship: $r_{XY} = r_{YX}$. The correlation between price and demand is algebraically and numerically identical to the correlation between demand and price.

Property V: Linear Specificity ($r = 0$ Does Not Mean Independence)

A value of $r = 0$ indicates the total absence of a linear relationship between $X$ and $Y$. However, it does not prove that the two variables are statistically independent. A perfect non-linear or curvilinear relationship (such as $Y = X^2$ over symmetric domains $[-3, 3]$) produces $r = 0$ even though $Y$ is completely deterministic from $X$.

2. Systematic Scale of Correlation Interpretation
Range of Correlation Coefficient ($r$) Standard Qualitative Interpretation Economic Implication
$$r = +1$$ Perfect Positive Correlation Variables move in exactly identical proportion in the same direction.
$$+0.75 \le r < +1$$ High Degree of Positive Correlation Very strong direct association; one variable closely tracks the other.
$$+0.25 \le r < +0.75$$ Moderate Degree of Positive Correlation Meaningful direct relationship, but influenced by other external factors.
$$0 < r < +0.25$$ Low Degree of Positive Correlation Weak positive association; poor predictive power.
$$r = 0$$ Zero / No Linear Correlation Complete absence of linear association between the series.
$$-0.25 < r < 0$$ Low Degree of Negative Correlation Weak inverse association.
$$-0.75 < r \le -0.25$$ Moderate Degree of Negative Correlation Substantial inverse relationship (e.g., standard price-demand elasticities).
$$-1 < r \le -0.75$$ High Degree of Negative Correlation Very strong inverse association.
$$r = -1$$ Perfect Negative Correlation Variables move in exactly identical proportion in opposite directions.
3. Statistical Significance: Probable Error & Standard Error

To evaluate whether a sample correlation coefficient $r$ reflects a genuine relationship in the underlying population or arose merely from random sampling fluctuations, statisticians utilize the Probable Error ($PE$):

$$PE_r = 0.6745 \times \frac{1 - r^2}{\sqrt{N}} \qquad \text{and} \qquad SE_r = \frac{1 - r^2}{\sqrt{N}}$$

Decision Criteria:

  • If $r < PE_r$, the correlation is statistically insignificant and should be discounted.
  • If $r > 6 \times PE_r$, the correlation is highly significant and points to a definitive relationship in the population.
  • The population correlation coefficient is expected to fall within limits: $r \pm PE_r$.

Spearman's Rank Correlation Coefficient (Ordinal Association)

1. Genesis & Purpose of Rank Correlation

In many economic and social inquiries, variables cannot be measured directly on a quantitative cardinal scale (e.g., intelligence, beauty, managerial competence, honesty, consumer preference, brand loyalty). However, statistical individuals can easily be arranged in serial order or ranked (1st, 2nd, 3rd, $\dots$).

To measure the association between two sets of qualitative rankings, the British psychologist and statistician Charles Edward Spearman developed the Rank Correlation Coefficient (denoted by $\rho$ or $R$) in 1904. It is a non-parametric metric that evaluates monotonic relationships without assuming a bivariate normal distribution.

2. Case I: When Ranks are Distinct (No Tied Ranks)

When each individual receives a unique, distinct rank in both series ($R_1$ and $R_2$), the formula is:

$$R = 1 - \frac{6 \sum D^2}{N(N^2 - 1)}$$

Where:

  • $D = R_1 - R_2$ is the difference between the ranks assigned to the same individual in variables $X$ and $Y$.
  • $\sum D^2$ is the sum of the squares of rank differences.
  • $N$ is the total number of paired observations / individuals ranked.
Automatic Computational Check:
The algebraic sum of rank differences must always be identically zero: $$\sum D = \sum (R_1 - R_2) \equiv 0$$ If $\sum D \neq 0$, an arithmetic error has occurred in assigning ranks or computing differences!
3. Case II: When Ranks are Repeated or Tied (Equal Values)

When two or more items have identical values in a series, they are assigned the average (arithmetic mean) of the ranks they would have occupied if they were slightly different. For example, if two students tie for the 3rd rank, they occupy positions 3 and 4; each is awarded rank $\frac{3 + 4}{2} = 3.5$. The next student receives rank 5.

Because tying reduces rank variance, a correction factor of $\frac{m(m^2 - 1)}{12}$ must be added to $\sum D^2$ for each tied group, where $m$ is the number of times a value repeats:

$$R = 1 - \frac{6 \left[ \sum D^2 + \sum \frac{m_i(m_i^2 - 1)}{12} \right]}{N(N^2 - 1)}$$

If value $X = 40$ repeats 3 times ($m_1 = 3$), its correction factor is $\frac{3(9 - 1)}{12} = \frac{24}{12} = 2.0$. If another value repeats twice ($m_2 = 2$), its correction factor is $\frac{2(4 - 1)}{12} = 0.5$.

4. Comparative Evaluation: Karl Pearson vs. Spearman's Rank
Comparative Feature Karl Pearson's Coefficient ($r$) Spearman's Rank Coefficient ($R$)
Nature of Data Cardinal quantitative metric data (measured in physical units: kg, ₹, metres). Ordinal ranking data or qualitative attributes (ranks, preferences, scores).
Underlying Distribution Assumes an underlying bivariate normal distribution. Non-parametric distribution-free; makes no normality assumptions.
Sensitivity to Outliers Highly sensitive to extreme values and outliers (since deviations are squared). Robust against extreme outliers because values are converted into bounded ranks.
Type of Association Measures strictly linear association. Measures monotonic association (whether strictly increasing or decreasing).
Computational Complexity Laborious for large datasets without computing software. Quick and simple, especially when ranks are already given.

Empirical Economic Applications, Limitations & Analytical Pitfalls

1. Real-World Applications in Applied Economics

Correlation analysis is embedded throughout empirical microeconomics, macroeconomic policymaking, and financial management:

  • Law of Demand & Elasticity: Economists test whether the empirical correlation between price and demand is consistently negative ($r < 0$). In commodity markets, high negative correlation confirms normal goods, while a positive correlation signals exceptional Giffen goods or Veblen goods.
  • Keynesian Consumption Function: John Maynard Keynes posited that consumption expenditure ($C$) is directly related to disposable income ($Y$). Empirical studies across Indian macroeconomic data reveal a high positive correlation ($r > +0.90$), validating the fundamental psychological law of consumption.
  • Phillips Curve & Monetary Trade-offs: The historical Phillips Curve examined the inverse correlation between the rate of unemployment and the rate of wage inflation ($r < 0$), providing central banks with critical policy insights regarding macroeconomic stabilization.
  • Modern Portfolio Diversification: In financial economics, Harry Markowitz demonstrated that constructing an investment portfolio of securities with negative or zero correlation minimizes overall investment variance (risk) without sacrificing expected returns.
2. Pitfalls, Traps & Spurious Correlation

In analyzing correlation, economists and students must remain vigilant against common analytical traps:

  1. Spurious / Nonsense Correlation: Two independent series may exhibit high correlation simply because both happen to share a common secular upward trend over time (e.g., both cumulative mobile phone sales in India and life expectancy have risen steadily over the past 30 years, but one does not cause the other).
  2. Third-Variable Distortion (Confounding Factors): Ice cream sales and drowning accidents display a strong positive correlation in summer months. Neither causes the other; both are driven by a confounding third variable: high ambient summer temperatures.
  3. Masked Non-Linearity: A low or zero Pearson coefficient ($r = 0$) does not mean variables are unrelated. The relationship between tax rates and total tax revenue (Laffer Curve) is inverted U-shaped: revenue rises with tax rates up to an optimal peak, then declines. Linear $r$ across the entire range would be close to zero, completely masking the profound non-linear relationship.
  4. Truncation of Range: Restricting the sample to a very narrow range (e.g., studying the correlation between intelligence and economic performance only among Nobel laureates) drastically reduces the observed correlation coefficient relative to the broader population.
3. Synthesis: Method Selection Flowchart

When confronted with an empirical problem, select the appropriate method using this diagnostic rule:

  • Quantitative cardinal data with linear relationship? $\rightarrow$ Use Karl Pearson's Direct or Short-cut Method ($r$).
  • Qualitative attributes (honesty, beauty, debate rankings) or ordinal data? $\rightarrow$ Use Spearman's Rank Correlation ($R$).
  • Initial exploratory visual inspection or non-linear trend detection? $\rightarrow$ Use Scatter Diagram.

Key Economic Identities, Formulas & Business Principles

Karl Pearson's Product-Moment Correlation Formulas
Spearman's Rank Correlation Formulas
Probable Error & Statistical Significance

Conceptual Solved Examples & Case Studies

Example 1
Step-by-Step Solution:
Step 1: Compute Arithmetic Means ($ar{X}$ and $ar{Y}$)
Number of pairs ($N$) = 5.
$$\sum X = 25 + 27 + 30 + 32 + 36 = 150 \implies ar{X} = rac{150}{5} = 30$$ $$\sum Y = 20 + 24 + 25 + 27 + 29 = 125 \implies ar{Y} = rac{125}{5} = 25$$ Both means are exact whole integers, so the Direct Method is ideal.

Step 2: Construct the Computation Table
Couple $X$ $x = X - 30$ $x^2$ $Y$ $y = Y - 25$ $y^2$ $xy$
125-52520-52525
227-3924-113
3300025000
432+2427+244
536+63629+41624
Total $\sum X = 150$ $\sum x = 0$ $\sum x^2 = 74$ $\sum Y = 125$ $\sum y = 0$ $\sum y^2 = 46$ $\sum xy = 56$
Step 3: Apply Pearson's Direct Formula $$r = rac{\sum xy}{\sqrt{\sum x^2 \cdot \sum y^2}} = rac{56}{\sqrt{74 \cdot 46}} = rac{56}{\sqrt{3404}} = rac{56}{58.3438} pprox +0.9598$$
Economic Interpretation:
The correlation coefficient between the age of husbands and wives is +0.96, indicating an exceptionally high degree of positive linear correlation.
Example 2
Step-by-Step Solution:
Step 1: Choose Assumed Means ($A_x, A_y$)
Let $A_x = 14$ and $A_y = 30$. Number of pairs $N = 6$.

Step 2: Construct the Short-Cut Table
Item $X$ $d_x = X - 14$ $d_x^2$ $Y$ $d_y = Y - 30$ $d_y^2$ $d_x d_y$
110-41640+10100-40
212-2438+864-16
3140032+240
416+2428-24-4
518+41625-525-20
620+63620-10100-60
Total - $\sum d_x = 6$ $\sum d_x^2 = 76$ - $\sum d_y = 3$ $\sum d_y^2 = 297$ $\sum d_x d_y = -140$
Step 3: Apply the Short-Cut Formula $$r = rac{N\sum d_x d_y - (\sum d_x)(\sum d_y)}{\sqrt{N\sum d_x^2 - (\sum d_x)^2} \cdot \sqrt{N\sum d_y^2 - (\sum d_y)^2}}$$ $$r = rac{6(-140) - (6)(3)}{\sqrt{6(76) - (6)^2} \cdot \sqrt{6(297) - (3)^2}}$$ $$r = rac{-840 - 18}{\sqrt{456 - 36} \cdot \sqrt{1782 - 9}} = rac{-858}{\sqrt{420} \cdot \sqrt{1773}}$$ $$r = rac{-858}{20.4939 \cdot 42.107} = rac{-858}{862.936} pprox -0.9943$$
Economic Interpretation:
The correlation coefficient is -0.99, proving an almost perfect negative linear correlation between price and demand, fully confirming the classical Law of Demand.
Example 3
Step-by-Step Solution:
Step 1: Set Assumed Means and Common Scale Factors
For $X$: Let $A_x = 300$, common scale factor $c_x = 100$. Thus $u = rac{X - 300}{100}$.
For $Y$: Let $A_y = 45$, common scale factor $c_y = 5$. Thus $v = rac{Y - 45}{5}$. Number of pairs $N = 5$.

Step 2: Construct the Step-Deviation Table
Unit $X$ $u = rac{X-300}{100}$ $u^2$ $Y$ $v = rac{Y-45}{5}$ $v^2$ $uv$
1100-2420-52510
2200-1135-242
33000045000
4400+1150+111
5500+2470+52510
Total - $\sum u = 0$ $\sum u^2 = 10$ - $\sum v = -1$ $\sum v^2 = 55$ $\sum uv = 23$
Step 3: Apply the Step-Deviation Formula $$r = rac{N\sum uv - (\sum u)(\sum v)}{\sqrt{N\sum u^2 - (\sum u)^2} \cdot \sqrt{N\sum v^2 - (\sum v)^2}}$$ $$r = rac{5(23) - (0)(-1)}{\sqrt{5(10) - (0)^2} \cdot \sqrt{5(55) - (-1)^2}} = rac{115}{\sqrt{50} \cdot \sqrt{275 - 1}} = rac{115}{\sqrt{50} \cdot \sqrt{274}}$$ $$r = rac{115}{7.071 \cdot 16.553} = rac{115}{117.046} pprox +0.9825$$
Economic Interpretation:
Because $r$ is invariant to scale factors $c_x$ and $c_y$, $r_{XY} = r_{uv} = \mathbf{+0.98}$, indicating an extremely high positive correlation between production scale and profitability.
Example 4
Step-by-Step Solution:
Step 1: Understand the Data
The ranks are already assigned directly by the judges ($R_1$ and $R_2$). Number of participants $N = 8$. All ranks are distinct integers from 1 to 8 without any ties.

Step 2: Construct the Difference Table
Participant Judge A ($R_1$) Judge B ($R_2$) $D = R_1 - R_2$ $D^2$
156-11
221+11
385+39
412-11
54400
667-11
738-525
873+416
Total - - $\sum D = 0$ (Verified) $\sum D^2 = 54$
Step 3: Apply Spearman's Distinct Rank Formula $$R = 1 - rac{6 \sum D^2}{N(N^2 - 1)}$$ $$R = 1 - rac{6(54)}{8(8^2 - 1)} = 1 - rac{324}{8(64 - 1)} = 1 - rac{324}{8(63)} = 1 - rac{324}{504}$$ $$R = 1 - 0.6429 = +0.3571 pprox +0.36$$
Interpretation:
The rank correlation coefficient is +0.36, indicating a moderate positive agreement between the aesthetic judgment of the two judges.
Example 5
Step-by-Step Solution:
Step 1: Assign Ranks with Average Rank Rule
In Economics ($X$): Values in descending order: 80 (rank 1), 75 (rank 2), 68 (rank 3), 64 and 64 (tied for ranks 4 and 5), 50 (rank 6).
For the tie at 64: Rank = $ rac{4 + 5}{2} = 4.5$. Here $m_1 = 2$.

In Statistics ($Y$): Values in descending order: 52 (rank 1), 48 (rank 2), 40 and 40 (tied for ranks 3 and 4), 35 (rank 5), 30 (rank 6).
For the tie at 40: Rank = $ rac{3 + 4}{2} = 3.5$. Here $m_2 = 2$.

Step 2: Construct the Difference Table
Student Economics ($X$) Rank $R_1$ Statistics ($Y$) Rank $R_2$ $D = R_1 - R_2$ $D^2$
1683403.5-0.50.25
2644.5403.5+1.01.00
3752482.00.00.00
4506306.00.00.00
5644.5355.0-0.50.25
6801521.00.00.00
Total - - - - $\sum D = 0$ $\sum D^2 = 1.50$
Step 3: Compute Tie Correction Factors Correction for $X$ ($m_1 = 2$): $ rac{2(2^2 - 1)}{12} = rac{2(3)}{12} = 0.5$.
Correction for $Y$ ($m_2 = 2$): $ rac{2(2^2 - 1)}{12} = rac{2(3)}{12} = 0.5$.
Total corrected term = $\sum D^2 + 0.5 + 0.5 = 1.50 + 1.0 = 2.50$.

Step 4: Apply the Tied Rank Formula $$R = 1 - rac{6 \left[ \sum D^2 + \sum rac{m_i(m_i^2 - 1)}{12} ight]}{N(N^2 - 1)} = 1 - rac{6(2.50)}{6(6^2 - 1)} = 1 - rac{15}{6(35)} = 1 - rac{15}{210} = 1 - 0.0714 = +0.9286$$
Interpretation:
The rank correlation coefficient is +0.93, indicating an exceptionally strong positive agreement between performance in Economics and Statistics.
Example 6
Step-by-Step Solution:

Part (a): Effect of Linear Transformations
We are given:

$$U = rac{X - 100}{-5} = - rac{1}{5}X + 20 \implies h = -5$$

$$V = rac{Y + 50}{10} = rac{1}{10}Y + 5 \implies k = +10$$

By Property II of Karl Pearson's coefficient:

$$r_{UV} = \left( rac{h \cdot k}{|h \cdot k|} ight) r_{XY}$$

Here $h \cdot k = (-5)(+10) = -50 < 0$. Because $h$ and $k$ have opposite signs, the sign of the correlation coefficient reverses:

$$r_{UV} = rac{-50}{|-50|} r_{XY} = -1 \cdot (+0.80) = -0.80$$


Part (b): Standard Error & Probable Error
Given $r = 0.80 \implies r^2 = 0.64$, and $N = 64 \implies \sqrt{N} = \sqrt{64} = 8$.

1. Standard Error ($SE_r$):

$$SE_r = rac{1 - r^2}{\sqrt{N}} = rac{1 - 0.64}{8} = rac{0.36}{8} = 0.045$$


2. Probable Error ($PE_r$):

$$PE_r = 0.6745 imes SE_r = 0.6745 imes 0.045 = 0.03035 pprox 0.030$$


3. Test of Statistical Significance:

$$6 imes PE_r = 6 imes 0.03035 = 0.1821$$

Because $r = 0.80$ is substantially greater than $6 imes PE_r$ ($0.80 > 0.1821$), the correlation coefficient is highly statistically significant, proving that the association is genuine and not an artifact of random sampling.

Common Misconceptions & Examiner Traps

Common Misconception

Confusing the square of the sum $(\sum d_x)^2$ with the sum of squares $\sum d_x^2$ in the short-cut formula.

Scientific Reality & Correction

$\sum d_x^2$ requires squaring each individual deviation first and then adding them together. $(\sum d_x)^2$ requires adding all deviations first and then squaring that single sum. These two quantities are completely different.

Common Misconception

Believing that a correlation of r = 0 proves that the two variables are completely independent.

Scientific Reality & Correction

A correlation of r = 0 only establishes the absence of a LINEAR relationship. Variables may share a strong non-linear functional relationship (e.g. Y = X²) while still producing r = 0.

Common Misconception

Forgetting to assign average ranks when values tie in Spearman's rank correlation.

Scientific Reality & Correction

When values tie (e.g., two items tied for ranks 3 and 4), you must assign the average rank (3.5) to both items, and the next item must receive rank 5 (not 4).

Visual Learning & Conceptual Map

r WBCHSE Class 11 Economics • Quantitative Statistics: Correlation Analysis Scatter Diagrams, Karl Pearson's Product-Moment Coefficient (r) & Spearman's Rank (R) 1. Scatter Diagrams & Direction of Association Perfect Positive (r = +1) Perfect Negative (r = -1) No Linear Assoc. (r = 0) Curvilinear (r = 0) Mathematical Properties: • Invariant to Change of Origin and Change of Scale • Pure dimensionless number strictly bounded: -1 ≤ r ≤ +1 • Symmetric with respect to variables: r(X,Y) = r(Y,X) 2. Karl Pearson's Coefficient of Correlation (r) Covariance to Product of Standard Deviations: r = Cov(X,Y) / (σx · σy) = Σ(x-x̄)(y-ȳ) / √[Σ(x-x̄)² · Σ(y-ȳ)²] Degree of Correlation Scale [-1 to +1]: -1.0 Perfect Neg (-1.0) -0.5 Mod Neg (-0.5) 0.0 Zero (0.0) +0.5 Mod Pos (+0.5) +1.0 Perfect Pos (+1.0) Direct: r = Σxy / √(Σx²·Σy²) Short-Cut: Assumed Mean dx, dy 3. Spearman's Rank Correlation & Economic Insights Distinct Ranks: R = 1 - [6·ΣD² / N(N²-1)] Tied Ranks: R = 1 - 6·[ΣD² + Σ(m³-m)/12] / N(N²-1) Computational Check: Sum of rank differences is zero (ΣD ≡ 0) Economic Applications & Caution: • Ideal for qualitative ordinal attributes: skill, beauty, competitive exam rankings • Price & Demand (Negative r) vs. Income & Consumption (Positive r) Core Principle: Correlation does NOT imply Causation (Spurious Correlation) TargetExams Academic Series • WBCHSE Quantitative Statistics • Bivariate Data & Correlation Analysis

Chapter Summary & 10 Key Takeaways

Takeaway 1
  1. Correlation is a bivariate statistical tool that discovers, measures, and interprets the degree and direction of linear association between two or more quantitative or qualitative variables.
Takeaway 2
  1. Correlation is classified by direction (Positive/Direct vs. Negative/Inverse), ratio of change (Linear vs. Non-Linear/Curvilinear), and number of variables (Simple, Multiple, Partial).
Takeaway 3
  1. Correlation establishes numerical association but does NOT prove causal determination. High correlation may result from direct causation, mutual interaction, a common third factor, or sheer historical coincidence (spurious correlation).
Takeaway 4
  1. The Scatter Diagram is a graphical representation of paired observations (x, y); it visually demonstrates the direction, dispersion, and curvature of relationship without calculating a numerical coefficient.
Takeaway 5
  1. Karl Pearson's Product-Moment Correlation Coefficient (r) is defined as the covariance of two variables divided by the product of their standard deviations: r = Cov(X, Y) / (σx · σy).
Takeaway 6
  1. Pearson's r is strictly bounded between -1 and +1 (-1 ≤ r ≤ +1). A value of +1 signifies perfect positive correlation, -1 signifies perfect negative correlation, and 0 denotes absence of linear correlation.
Takeaway 7
  1. Pearson's r is completely independent of the change of origin and change of scale (magnitude), provided the scaling constants have the same sign; opposite signs reverse the algebraic sign of r.
Takeaway 8
  1. Spearman's Rank Correlation Coefficient (R) is a non-parametric technique used when data are in the form of ordinal ranks or qualitative attributes (beauty, skill, preference): R = 1 - [6ΣD² / N(N² - 1)].
Takeaway 9
  1. When ranks are tied, an average rank is assigned to equal items, and a correction factor of m(m² - 1)/12 is added to ΣD² for each group of m repeating items.
Takeaway 10
  1. In statistical hypothesis testing, correlation r is deemed significant if r > 6 × PE_r, where Probable Error is PE_r = 0.6745(1 - r²) / √N.

Check Your Understanding (Diagnostic Practice Questions)

Diagnostic questions testing core conceptual clarity. Answers are hidden initially — solve each problem first, then click to reveal the step-by-step verified solution.

1
Reveal Answer & Explanation
Answer:
2
Reveal Answer & Explanation
Answer:
3
Reveal Answer & Explanation
Answer:
4
Reveal Answer & Explanation
Answer:
5
Reveal Answer & Explanation
Answer:
Finished Studying This Chapter?
READY TO PRACTICE?

Timed CBT Practice Tests (Exam Simulator)

Put your concepts to the test with official curriculum-aligned Foundation and Advanced practice tests. Get instant accuracy scores, time metrics, and step-by-step verified explanations.