← Back to Pandas Course | Chapter 11: File I/O | Lesson 5 of 7

read_html का उपयोग

read_html किसी web page (या HTML string) पर हर table ढूंढता है और उन्हें DataFrames के रूप में return करता है।

In this page:

  1. read_html
Syntax
python
tables = pd.read_html("page.html")
df = tables[0]

read_html

pd.read_html DataFrames की एक list return करता है, हर table के लिए एक। इसे lxml parser (या html5lib के साथ BeautifulSoup) चाहिए। एक URL, file या StringIO में wrapped HTML string पास करें। किसी specific text वाला table चुनने के लिए match का उपयोग करें।

Note: list index से अपनी ज़रूरत का table चुनें, जैसे tables[0]।

उदाहरण: read_html

python
import io
import lxml
import pandas as pd

html = "<table><tr><th>city</th><th>pop</th></tr><tr><td>Oslo</td><td>700</td></tr><tr><td>Rome</td><td>2800</td></tr></table>"
tables = pd.read_html(io.StringIO(html))
print(len(tables))
print(tables[0])

# Output:
# 1
#    city   pop
# 0  Oslo   700
# 1  Rome  2800
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the view already swaps in the hi_ translation fields (or blanks these out if untranslated), so this renders correctly for both languages without a lang_code check here. #}
आम गलतियां
  1. यह भूल जाना कि यह एक list return करता है
  2. lxml या html5lib का न होना
  3. नए pandas में सीधे raw HTML string पास करना
चैप्टर सारांश
  • read_html DataFrames की list return करता है
  • इसे lxml या bs4 plus html5lib चाहिए
  • strings को StringIO में wrap करें
  • match tables को filter करता है
🔒

Chapter Quiz — Complete all 7 topics to unlock

0/7 topics done

Complete these topics first:

Login to run this code

C/C++/Java/PHP execution requires a free account. Your code is saved — you'll land right back in the editor after logging in.