Categorical Dtype की जानकारी
category dtype दोहराए गए text को एक बार store करता है और उसकी ओर point करता है, जिससे memory बचती है और grouping तेज़ होती है।
In this page:
Syntax
df["column"] = df["column"].astype("category")
df["column"].cat.categories
Categorical dtype
astype("category") कुछ distinct values वाले column को integer codes plus एक lookup table में बदल देता है। यह object strings से कहीं कम memory का उपयोग करता है और एक custom order रख सकता है। Ordered categoricals logically sort और compare होते हैं।
Note:
Categoricals country, status या size जैसे कई repeats वाले columns के लिए शानदार हैं।
उदाहरण: Categorical dtype
import pandas as pd
s = pd.Series(["low", "high", "medium", "low"] * 1000)
c = s.astype("category")
print(s.memory_usage(deep=True) > c.memory_usage(deep=True))
sizes = pd.Categorical(["M", "S", "L"], categories=["S", "M", "L"], ordered=True)
print(sizes.sort_values().tolist())
# Output:
# True
# ['S', 'M', 'L']
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the
view already swaps in the hi_ translation fields (or blanks these
out if untranslated), so this renders correctly for both languages
without a lang_code check here. #}
आम गलतियां
- लगभग unique values के लिए category का उपयोग करना
- categories को extend किए बिना unseen values जोड़ना
- sorting के लिए order भूल जाना
चैप्टर सारांश
- category repeats को codes के रूप में store करता है
- Memory बचाता है
- custom ordering को support करता है
- low-cardinality columns के लिए सबसे अच्छा
🔒
Chapter Quiz — Complete all 6 topics to unlock
0/6 topics done
Complete these topics first: