Duplicate Rows को संभालना
Duplicated rows आपके numbers inflate कर सकती हैं, इसलिए इन्हें spot और drop करना सीखें।
In this page:
Syntax
df.duplicated()
df = df.drop_duplicates()
Duplicate Rows
duplicated() repeated rows flag करता है (first occurrence default से False है)। drop_duplicates() इन्हें remove करता है, और subset check को कुछ columns तक limit करता है। keep="last" या keep=False बदलते हैं कि कौन-सी copies retained रहती हैं।
Note:
df.duplicated().sum() duplicates count करता है।
उदाहरण: Duplicate rows
import pandas as pd
df = pd.DataFrame({"id": [1, 2, 2, 3], "name": ["Ann", "Bob", "Bob", "Cy"]})
print(df.duplicated().tolist())
print(df.drop_duplicates())
print(df.drop_duplicates(subset="name", keep="last"))
# Output:
# [False, False, True, False]
# id name
# 0 1 Ann
# 1 2 Bob
# 3 3 Cy
# id name
# 0 1 Ann
# 2 2 Bob
# 3 3 Cy
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the
view already swaps in the hi_ translation fields (or blanks these
out if untranslated), so this renders correctly for both languages
without a lang_code check here. #}
आम गलतियां
- यह decide किए बिना duplicates drop करना कि कौन-से columns एक duplicate define करते हैं
- यह भूल जाना कि first occurrence रखी जाती है
- drop करने के बाद index reset न करना
चैप्टर सारांश
- duplicated repeats flag करता है
- drop_duplicates इन्हें remove करता है
- subset key columns pick करता है
- keep चुनता है कौन-सी copy रहती है
🔒
Chapter Quiz — Complete all 7 topics to unlock
0/7 topics done
Complete these topics first: