Duplicate Columns को संभालना
जब दोनों tables column names share करते हैं, तो merge के बाद suffixes उन्हें अलग रखते हैं।
In this page:
Syntax
merged = pd.merge(left, right, on="key_column", suffixes=("_left", "_right"))
Handling duplicate columns
दोनों tables में मौजूद non-key columns को default रूप से _x और _y suffixes मिलते हैं। स्पष्ट नामों के लिए suffixes=("_left", "_right") पास करें। टकराव से बचने के लिए merge करने से पहले columns को drop या rename भी किया जा सकता है।
Note:
rows का गुणा होने से पहले duplicate keys पकड़ने के लिए validate का उपयोग करें।
उदाहरण: Handling duplicate columns
import pandas as pd
a = pd.DataFrame({"id": [1, 2], "score": [10, 20]})
b = pd.DataFrame({"id": [1, 2], "score": [15, 25]})
print(a.merge(b, on="id"))
print(a.merge(b, on="id", suffixes=("_2023", "_2024")))
# Output:
# id score_x score_y
# 0 1 10 15
# 1 2 20 25
# id score_2023 score_2024
# 0 1 10 15
# 1 2 20 25
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the
view already swaps in the hi_ translation fields (or blanks these
out if untranslated), so this renders correctly for both languages
without a lang_code check here. #}
आम गलतियां
- _x और _y को बिना समझाए छोड़ देना
- columns के गलत subset पर merge करना
- duplicate keys की जांच न करना
चैप्टर सारांश
- Overlapping columns को suffixes मिलते हैं
- Custom suffixes ज़्यादा स्पष्ट होते हैं
- टकराव से बचने के लिए merge से पहले rename करें
- validate key समस्याओं को पकड़ता है
🔒
Chapter Quiz — Complete all 7 topics to unlock
0/7 topics done
Complete these topics first: