Raw data is almost never usable as-is. It often contains errors, inconsistencies, or missing information.
If you skip cleaning, your analysis can become completely misleading.
🔴 Example:
Imagine a dataset of salaries:
- ₹50,000
- ₹60,000
- ₹5,000,000 (wrong entry)
That one incorrect value will inflate the average and distort your conclusions.
⚠️ Common Data Problems:
1. Missing Values
- Empty cells or
NaN - Example: Age column has blanks
2. Duplicate Entries
- Same record appears multiple times
- Example: Same customer listed twice
3. Incorrect Formats
- Numbers stored as text
- Dates in inconsistent formats (e.g.,
12-01-24vs2024/01/12)
💡 Key Idea:
Clean data = reliable insights
Dirty data = wrong decisions