← Back to Pandas Course | Chapter 11: File I/O | Lesson 6 of 7

बड़ी Files को Chunk करना

किसी बड़ी file को chunks में पढ़ना एक समय में एक slice process करके memory कम रखता है।

In this page:

  1. Chunking large files
Syntax
python
for chunk in pd.read_csv("file.csv", chunksize=n):
    # process chunk (a DataFrame)

Chunking large files

read_csv(chunksize=n) हर एक में n rows वाले DataFrames का एक iterator return करता है। हर chunk को process करें, जैसे aggregate करना, और छोटे results को combine करें। इससे आप ऐसी files handle कर सकते हैं जो memory में fit नहीं होतीं।

Note: chunking से पहले memory कम करने के लिए usecols और dtype पर भी विचार करें।

उदाहरण: Chunking large files

python
import io
import pandas as pd

csv = "x\n" + "\n".join(str(i) for i in range(1, 11))
total = 0
for chunk in pd.read_csv(io.StringIO(csv), chunksize=4):
    total += chunk["x"].sum()
    print("chunk rows:", len(chunk))
print("total:", total)

# Output:
# chunk rows: 4
# chunk rows: 4
# chunk rows: 2
# total: 55
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the view already swaps in the hi_ translation fields (or blanks these out if untranslated), so this renders correctly for both languages without a lang_code check here. #}
आम गलतियां
  1. सभी chunks को concatenate करके फायदा खो देना
  2. यह भूल जाना कि iterator एक बार उपयोग होता है
  3. chunks के बीच गलत तरीके से aggregate करना
चैप्टर सारांश
  • chunksize एक iterator return करता है
  • हर chunk को अलग से process करें
  • छोटे results को combine करें
  • usecols और dtype memory घटाते हैं
🔒

Chapter Quiz — Complete all 7 topics to unlock

0/7 topics done

Complete these topics first:

Login to run this code

C/C++/Java/PHP execution requires a free account. Your code is saved — you'll land right back in the editor after logging in.