← Back to Pandas Course | Chapter 5: Data Cleaning | Lesson 2 of 7

Duplicate Rows को संभालना

Duplicated rows आपके numbers inflate कर सकती हैं, इसलिए इन्हें spot और drop करना सीखें।

In this page:

  1. Duplicate Rows
Syntax
python
df.duplicated()
df = df.drop_duplicates()

Duplicate Rows

duplicated() repeated rows flag करता है (first occurrence default से False है)। drop_duplicates() इन्हें remove करता है, और subset check को कुछ columns तक limit करता है। keep="last" या keep=False बदलते हैं कि कौन-सी copies retained रहती हैं।

Note: df.duplicated().sum() duplicates count करता है।

उदाहरण: Duplicate rows

python
import pandas as pd

df = pd.DataFrame({"id": [1, 2, 2, 3], "name": ["Ann", "Bob", "Bob", "Cy"]})
print(df.duplicated().tolist())
print(df.drop_duplicates())
print(df.drop_duplicates(subset="name", keep="last"))

# Output:
# [False, False, True, False]
#    id name
# 0   1  Ann
# 1   2  Bob
# 3   3   Cy
#    id name
# 0   1  Ann
# 2   2  Bob
# 3   3   Cy
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the view already swaps in the hi_ translation fields (or blanks these out if untranslated), so this renders correctly for both languages without a lang_code check here. #}
आम गलतियां
  1. यह decide किए बिना duplicates drop करना कि कौन-से columns एक duplicate define करते हैं
  2. यह भूल जाना कि first occurrence रखी जाती है
  3. drop करने के बाद index reset न करना
चैप्टर सारांश
  • duplicated repeats flag करता है
  • drop_duplicates इन्हें remove करता है
  • subset key columns pick करता है
  • keep चुनता है कौन-सी copy रहती है
🔒

Chapter Quiz — Complete all 7 topics to unlock

0/7 topics done

Complete these topics first:

Login to run this code

C/C++/Java/PHP execution requires a free account. Your code is saved — you'll land right back in the editor after logging in.