Statistical data collected during an economic survey or census investigation are known as Raw Data (or crude data). Raw data are an unstructured, chaotic collection of observations recorded in the order of their enumeration. In this unorganized form, they possess three major deficiencies: they fail to convey any immediate economic meaning, occupy immense space, and cannot be utilized for comparative analysis or algebraic computations. Organisation of data refers to the systematic arrangement and condensation of raw data into orderly classes, groups, and tables to bring out their underlying relationships and salient features.
Classification is the primary operational step in data organisation. It is defined as:
"Classification is the process of arranging data into sequences and groups according to their common characteristics or separating them into different but related parts." — Horace Secrist
The principal objectives of classification in economic statistics include:
- Simplification and Condensation: Reducing immense, unwieldy mass data into compact, homogeneous groups that human intelligence can easily absorb.
- Highlighting Similarities and Contrasts: Segregating data points possessing common attributes from those with distinct traits (e.g., separating literate agricultural labourers from illiterate ones).
- Facilitating Meaningful Comparison: Enabling comparative evaluations between demographic groups, time periods, or geographic regions.
- Preparing Data for Tabulation: Serving as the indispensable preparatory bridge between raw collection and formal statistical tabulation.
- Exhaustiveness / Comprehensiveness: Every single item in the raw dataset must find a place in one of the classes without any omission.
- Mutual Exclusivity: Classes must be non-overlapping; no single observation should qualify for more than one class.
- Stability: The basis of classification must remain uniform throughout the entire investigation to prevent bias.
- Flexibility: The classification scheme should be adaptable to changing circumstances without destroying comparative validity.
- Homogeneity: All units placed within a given class must share fundamentally similar characteristics.
| Classification Basis | Defining Principle | Economic Example |
|---|---|---|
| 1. Chronological (Temporal) | Observations are arranged sequentially with respect to time (years, quarters, months, or decades). | India's decadal population census figures from 1951 to 2021; annual foodgrain production of West Bengal. |
| 2. Geographical (Spatial) | Data are grouped according to geographical or spatial locations (countries, states, districts, rural/urban). | State-wise per capita income across Indian states; district-wise jute production in Bengal. |
| 3. Qualitative | Data are classified on the basis of non-measurable descriptive attributes (gender, religion, literacy, employment status).
|
Classification of workforce by gender, technical skill levels, and rural-urban domicile. |
| 4. Quantitative | Observations are classified on the basis of quantifiable, numerically measurable economic variables (income, expenditure, output, height, weight). | Distribution of industrial workers grouped into wage brackets (e.g., Rs. 10,000–15,000, Rs. 15,000–20,000). |
In quantitative classification, a measurable phenomenon that varies in magnitude across observation units is termed a variable:
- Discrete Variable: A variable that increases by distinct, finite, disconnected jumps and assumes only isolated exact integers. Fractional values are physically impossible (e.g., number of children per household, number of road accidents per month, number of printing errors per page).
- Continuous Variable: A variable capable of assuming any real numerical value (including infinite fractions and decimals) within a continuous numerical interval (e.g., worker wages, body weight, temperature, monthly electricity consumption, agricultural yield per acre).