Raw data is almost never usable as-is. It often contains errors, inconsistencies, or missing information.

If you skip cleaning, your analysis can become completely misleading.

🔴 Example:

Imagine a dataset of salaries:

  • ₹50,000
  • ₹60,000
  • ₹5,000,000 (wrong entry)

That one incorrect value will inflate the average and distort your conclusions.

⚠️ Common Data Problems:

1. Missing Values

  • Empty cells or NaN
  • Example: Age column has blanks

2. Duplicate Entries

  • Same record appears multiple times
  • Example: Same customer listed twice

3. Incorrect Formats

  • Numbers stored as text
  • Dates in inconsistent formats (e.g., 12-01-24 vs 2024/01/12)

💡 Key Idea:

Clean data = reliable insights
Dirty data = wrong decisions