read_html का उपयोग
read_html किसी web page (या HTML string) पर हर table ढूंढता है और उन्हें DataFrames के रूप में return करता है।
In this page:
Syntax
tables = pd.read_html("page.html")
df = tables[0]
read_html
pd.read_html DataFrames की एक list return करता है, हर table के लिए एक। इसे lxml parser (या html5lib के साथ BeautifulSoup) चाहिए। एक URL, file या StringIO में wrapped HTML string पास करें। किसी specific text वाला table चुनने के लिए match का उपयोग करें।
Note:
list index से अपनी ज़रूरत का table चुनें, जैसे tables[0]।
उदाहरण: read_html
import io
import lxml
import pandas as pd
html = "<table><tr><th>city</th><th>pop</th></tr><tr><td>Oslo</td><td>700</td></tr><tr><td>Rome</td><td>2800</td></tr></table>"
tables = pd.read_html(io.StringIO(html))
print(len(tables))
print(tables[0])
# Output:
# 1
# city pop
# 0 Oslo 700
# 1 Rome 2800
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the
view already swaps in the hi_ translation fields (or blanks these
out if untranslated), so this renders correctly for both languages
without a lang_code check here. #}
आम गलतियां
- यह भूल जाना कि यह एक list return करता है
- lxml या html5lib का न होना
- नए pandas में सीधे raw HTML string पास करना
चैप्टर सारांश
- read_html DataFrames की list return करता है
- इसे lxml या bs4 plus html5lib चाहिए
- strings को StringIO में wrap करें
- match tables को filter करता है
🔒
Chapter Quiz — Complete all 7 topics to unlock
0/7 topics done
Complete these topics first: