बड़ी Files को Chunk करना
किसी बड़ी file को chunks में पढ़ना एक समय में एक slice process करके memory कम रखता है।
In this page:
Syntax
for chunk in pd.read_csv("file.csv", chunksize=n):
# process chunk (a DataFrame)
Chunking large files
read_csv(chunksize=n) हर एक में n rows वाले DataFrames का एक iterator return करता है। हर chunk को process करें, जैसे aggregate करना, और छोटे results को combine करें। इससे आप ऐसी files handle कर सकते हैं जो memory में fit नहीं होतीं।
Note:
chunking से पहले memory कम करने के लिए usecols और dtype पर भी विचार करें।
उदाहरण: Chunking large files
import io
import pandas as pd
csv = "x\n" + "\n".join(str(i) for i in range(1, 11))
total = 0
for chunk in pd.read_csv(io.StringIO(csv), chunksize=4):
total += chunk["x"].sum()
print("chunk rows:", len(chunk))
print("total:", total)
# Output:
# chunk rows: 4
# chunk rows: 4
# chunk rows: 2
# total: 55
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the
view already swaps in the hi_ translation fields (or blanks these
out if untranslated), so this renders correctly for both languages
without a lang_code check here. #}
आम गलतियां
- सभी chunks को concatenate करके फायदा खो देना
- यह भूल जाना कि iterator एक बार उपयोग होता है
- chunks के बीच गलत तरीके से aggregate करना
चैप्टर सारांश
- chunksize एक iterator return करता है
- हर chunk को अलग से process करें
- छोटे results को combine करें
- usecols और dtype memory घटाते हैं
🔒
Chapter Quiz — Complete all 7 topics to unlock
0/7 topics done
Complete these topics first: