NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

· 19 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

Updated 2026-09-17

Re-tested on pandas 3.0.5, where pd.read_html("<table>…") with a raw string no longer works. Every option on this page now runs against one fixture page with eleven tables in the evidence packet, with the DataFrame output shown. The old page's df.append, chunksize= and usecols= recipes, which do not run on pandas 2.0 or later, are shown failing and replaced.

pandas.read_html() finds every <table> in a page and returns a list of DataFrames. It takes a file path, a URL, or a file-like object (StringIO around HTML you already have; a raw string no longer works on pandas 3). The three-line version, on a file:

import pandas as pd
PAGE = "fixtures/tables.html"
tables = pd.read_html(PAGE)
print("file:", len(tables), "tables; shapes:", [df.shape for df in tables])
print(tables[0])
file: 12 tables; shapes: [(3, 3), (3, 4), (2, 5), (3, 3), (2, 3), (2, 3), (2, 2), (2, 2), (5, 2), (2, 2), (3, 2), (3, 3)]
Name Age City
0 Ada 36 London
1 Grace 45 Arlington
2 Linus 28 Helsinki

The same call on HTML you already hold, and on a URL:

pd.read_html(StringIO(html)): 12
pd.read_html(open(PAGE, 'rb')): 12
pd.read_html(URL): 12

The rest of this page is what happens after that line: how to get the one table you want, why headers, numbers and links come out the way they do, which parser is running, why a Wikipedia URL returns 403, and which idioms from older tutorials (including the 2024 version of this page) raise on pandas 3.

Video Tutorial​

Versions and the fixture​

python: Python 3.12.11
pandas 3.0.5
numpy 2.5.3
lxml 6.1.3
beautifulsoup4 4.15.0
html5lib 1.1
requests 2.34.2
libxml2 2.14.6

read_html needs a parser: lxml, or beautifulsoup4 plus html5lib. Install all three, as the IO guide recommends, because the default flavor tries lxml first and then bs4 on any ValueError, including "no table matched", and without bs4 that fallback is an ImportError (shown below):

pip install pandas lxml beautifulsoup4 html5lib

The fixture is one HTML page with eleven tables (fixtures/tables.html in the packet, also served at https://scrapingant.github.io/scrapingant-examples/fixtures/html-tables.html for the URL examples), each written to trigger one behaviour: a plain table, a <thead> with thousands separators and currency, a two-row header with colspan, rowspan/colspan in the body, codes with leading zeros, dates, links, a display:none row, a nested table, a table whose first row is <td> cells, and European number formatting. Every Python block below is a slice of a packet script and every output block is the captured file; the packet's show() helper prints a value, or the exception type and the first line of its message, so the FileNotFoundError lines below are truncated (the real message contains the whole HTML). The snippets assume:

import re
from io import StringIO
import pandas as pd
import requests
from bs4 import BeautifulSoup
PAGE = "fixtures/tables.html"
URL = "https://scrapingant.github.io/scrapingant-examples/fixtures/html-tables.html"
WIKI = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population"
UA = {"User-Agent": "Mozilla/5.0 (compatible; scrapingant-examples/1.0; +https://github.com/ScrapingAnt/scrapingant-examples)"}

and one helper that returns the table with a given id:

def table(attrs_id, **kw):
"""The one table with id=attrs_id from the fixture page."""
return pd.read_html(PAGE, attrs={"id": attrs_id}, **kw)[0]

Input: file, string, URL​

read_html accepts a path, a URL, or a file-like object. It does not accept the HTML itself as a string any more: pandas 2.1 deprecated that and 3.0 removed it, so a raw string is treated as a path and you get FileNotFoundError with your HTML in the message. Wrap the string in io.StringIO (or BytesIO for bytes):

pd.read_html(html) with a str (pandas 3): FileNotFoundError: [Errno 2] No such file or directory: <!DOCTYPE html>
pd.read_html(StringIO(html)): 12
pd.read_html(open(PAGE, 'rb')): 12
pd.read_html(URL): 12
pd.read_html('https://' + ...)[1]:
Product Price Units sold Share
0 Ant Farm Deluxe $1,234.50 12345 45.5%
1 Ant Farm Mini $99.00 1020 3.8%
2 Magnifier $12.25 987654 50.7%

The fixture has eleven <table> elements but the list has twelve entries, because the nested table inside table 9 counts as its own table (more on that below). The return value is always a list, even for one table; index it or use attrs to get a single DataFrame.

Picking the table you want​

Two filters, both applied before parsing into a DataFrame: match (a string or compiled regex tested against the text nodes inside each table; the default .+ matches every non-empty table) and attrs (HTML attributes of the <table> element). They combine.

match='Ant Farm': [(3, 4), (2, 2)]
match='Ant Farm Deluxe': [(3, 4), (2, 2)]
match=re.compile(r'^Q[1-4]$'): [(2, 5)]
attrs={'id': 'prices'}:
Product Price Units sold Share
0 Ant Farm Deluxe $1,234.50 12345 45.5%
1 Ant Farm Mini $99.00 1020 3.8%
2 Magnifier $12.25 987654 50.7%
attrs={'class': 'data'}: [(3, 4)]
match='Magnifier', attrs={'class': 'data'}: [(3, 4)]
match='Farm Deluxe' (part of one cell): [(3, 4), (2, 2)]
match='Ant Farm Deluxe manual' (spans two cells): ValueError: No tables found matching pattern 'Ant Farm Deluxe manual'
match='no such text': ValueError: No tables found matching pattern 'no such text'
match='no such text', flavor='lxml' (lxml's own message): ValueError: No tables found matching regex 'no such text'
match='no such text' with only lxml installed (bs4 blocked): ImportError: `Import beautifulsoup4` failed. Use pip or conda to install the beautifulsoup4 package.
attrs={'id': 'nope'}: ValueError: No tables found
pd.read_html('fixtures/no-tables.html'): ValueError: No tables found
header-only table (empty tbody, as a JS page skeleton): (0, 2)
its columns: ['A', 'B']

match is a regex search against each text node inside the table, in practice each cell: "Ant Farm" also returns the links table, which has a cell containing "Ant Farm Deluxe"; the anchored ^Q[1-4]$ matches a whole cell; and "Ant Farm Deluxe manual" finds nothing although both strings sit in adjacent cells. attrs={"id": …} is the precise tool when the page gives the table an id. When nothing matches, read_html raises ValueError rather than returning an empty list (the reference allows rare exceptions), so a page that lost its table breaks loudly. Notice the two messages: "matching regex" is lxml's, "matching pattern" is bs4's. With the default flavor, lxml's no-match ValueError triggers the bs4 fallback, so the page is parsed a second time with html5lib and the error you see is bs4's; with only lxml installed that fallback fails with the ImportError shown, unless you pass flavor="lxml". A table that is present but empty, such as the header-only skeleton a JavaScript page ships before its script runs, is not an error: you get a DataFrame with the column names and no rows.

Headers​

By default, rows inside <thead> (or, when there is no <thead>, the top rows made only of <th> cells) become the column index. Two header rows become a MultiIndex, and colspan in the header is expanded so that every column has a label at every level:

df = table("two-row-header")
two-row thead -> MultiIndex columns:
Region 2025 2026
Region Q1 Q2 Q1 Q2
0 North 10 12 14 16
1 South 7 8 9 11
df.columns: [('Region', 'Region'), ('2025', 'Q1'), ('2025', 'Q2'), ('2026', 'Q1'), ('2026', 'Q2')]

The rowspan="2" header cell is repeated at both levels. To get flat names, join the levels (the repeated Region_Region is the cost of this one-liner; rename it afterwards):

df.columns = ["_".join(str(x) for x in col if str(x) != "nan") for col in df.columns]
flattened:
Region_Region 2025_Q1 2025_Q2 2026_Q1 2026_Q2
0 North 10 12 14 16
1 South 7 8 9 11
header=1 (second thead row only):
Region Q1 Q2 Q1.1 Q2.1
0 North 10 12 14 16
1 South 7 8 9 11

header=1 keeps only the second header row and pandas de-duplicates the repeated Q1/Q2 with .1. Spans in the body are expanded the same way (the rowspan value "Tools" is repeated, the colspan value "sold out" fills both cells), and a table whose first row is <td> cells gets integer column names until you say header=0:

spans (rowspan/colspan expanded):
Group Item Qty
0 Tools Tweezers 4
1 Tools Brush 2
2 Food sold out sold out
no thead, td first row: default:
0 1
0 Metric Value
1 Uptime 99.9
2 Errors 3
no thead, td first row: header=0:
Metric Value
0 Uptime 99.9
1 Errors 3.0

header=None is the default and means "infer", so on a table with a <th> row it changes nothing; there is no option that keeps a <th> row as data (the 2024 page said header=None meant "no header row"; it does not). index_col=0 moves a column into the index; skiprows=1 drops the first body row and skiprows=[1] drops row 1 specifically:

simple: header=None:
Name Age City
0 Ada 36 London
1 Grace 45 Arlington
2 Linus 28 Helsinki
simple: index_col=0:
Age City
Name
Ada 36 London
Grace 45 Arlington
Linus 28 Helsinki
simple: skiprows=1:
Ada 36 London
0 Grace 45 Arlington
1 Linus 28 Helsinki
simple: skiprows=[1]:
Name Age City
0 Grace 45 Arlington
1 Linus 28 Helsinki
prices (has thead): skiprows=1:
Ant Farm Deluxe $1,234.50 12,345 45.5%
0 Ant Farm Mini $99.00 1020 3.8%
1 Magnifier $12.25 987654 50.7%

Note what skiprows=1 did: the header row counts as a row, so skipping one row skipped the header and promoted "Ada" to a column name, and the prices table with a proper <thead> lost its header the same way. The reference says it: header is applied after skiprows. Use skiprows=[n] with the body row numbers you mean, or drop rows from the DataFrame afterwards.

Numbers, currency, leading zeros, dates​

thousands="," is the default, so 12,345 becomes the integer 12345 while $1,234.50 and 45.5% stay strings (pandas 3 shows the dtype as str). Clean those with converters: a converter receives each cell's text and replaces type inference for that column, so the column's dtype is whatever the function returns:

clean = {"Price": lambda s: float(s.replace("$", "").replace(",", "")), "Share": lambda s: float(s.rstrip("%")) / 100}
prices: dtypes:
Product str
Price str
Units sold int64
Share str
dtype: object
prices: thousands=None:
Product str
Price str
Units sold str
Share str
dtype: object
prices: converters for $ and %:
Product Price Units sold Share
0 Ant Farm Deluxe 1234.50 12345 0.455
1 Ant Farm Mini 99.00 1020 0.038
2 Magnifier 12.25 987654 0.507

Leading zeros are lost by the same type inference: 007 becomes 7 and the postcode 01234 becomes 1234. converters={"Code": str} keeps the text, and an identity converter shows why: the column stays str only because inference was switched off for it:

codes: default (leading zeros lost):
Code Postcode Label
0 7 1234 agent
1 42 501 answer
codes: converters={'Code': str, 'Postcode': str}:
Code Postcode Label
0 007 01234 agent
1 042 00501 answer
codes: converters={'Code': lambda s: s} (identity) dtypes:
Code str
Postcode int64
Label str
dtype: object

European formatting is the trap that produces wrong numbers silently rather than strings: with the default thousands=",", the growth figure 3,5 is read as 35 and -1,2 as -12. Set thousands="." and decimal=",":

european: default:
Country Revenue Growth
0 Germany 1.234.567,89 35
1 France 987.654,32 -12
2 Spain NaN 0
european: thousands='.', decimal=',':
Country Revenue Growth
0 Germany 1234567.89 3.5
1 France 987654.32 -1.2
2 Spain NaN 0.0

n/a became NaN without any option, because it is in pandas' default NA list; keep_default_na=False alone keeps it as the string n/a, and na_values adds to the list (or replaces it when combined with keep_default_na=False):

european: keep_default_na=False alone:
Country Revenue Growth
0 Germany 1234567.89 3.5
1 France 987654.32 -1.2
2 Spain n/a 0.0
european: keep_default_na=False, na_values=['n/a']:
Country Revenue Growth
0 Germany 1234567.89 3.5
1 France 987654.32 -1.2
2 Spain NaN 0.0

Dates follow read_csv semantics: parse_dates=True parses the index and nothing else, so on a table with no index_col it changes nothing; name the column (parse_dates=["Shipped"]), or set index_col=0 and let True parse the index, or convert after reading with pd.to_datetime. dtype_backend="numpy_nullable" gives string and Int64 instead of str and int64:

dates: default dtypes:
Shipped str
Delivered str
Parcels int64
dtype: object
dates: parse_dates=True:
Shipped str
Delivered str
Parcels int64
dtype: object
dates: parse_dates=['Shipped']:
Shipped datetime64[us]
Delivered str
Parcels int64
dtype: object
dates: index_col=0, parse_dates=True (index parsed): DatetimeIndex(['2026-09-01', '2026-09-10'], dtype='datetime64[us]', name='Shipped', freq=None)
dates: pd.to_datetime after reading:
Shipped datetime64[us]
Delivered str
Parcels int64
dtype: object
dates: dtype_backend='numpy_nullable':
Shipped string
Delivered string
Parcels Int64
dtype: object

By default the <a> text is kept and the href is dropped. extract_links="body" (or "header", "footer", "all") turns every cell in that section into a (text, href) tuple, with None where there was no link. With "all" the column labels become tuples too, so you index with the tuple:

links: default:
Product Docs
0 Ant Farm Deluxe manual
1 Magnifier none
links: extract_links='body':
Product Docs
0 (Ant Farm Deluxe, None) (manual, https://example.com/deluxe)
1 (Magnifier, None) (none, None)
links: extract_links='all':
(Product, /docs/product) (Docs, None)
0 (Ant Farm Deluxe, None) (manual, https://example.com/deluxe)
1 (Magnifier, None) (none, None)
columns: [('Product', '/docs/product'), ('Docs', None)]
hrefs of the Docs column: ['https://example.com/deluxe', None]

Hidden rows and nested tables​

displayed_only=True (the default) drops elements whose inline style attribute contains display:none, and nothing else. Rows hidden through a class and a stylesheet, the hidden attribute, visibility:hidden, or even upper-case DISPLAY: NONE are kept, with both parsers. Turn it off when the hidden row is the data:

hidden: default (displayed_only=True):
Sku Stock
0 AF-1 5
1 AF-3 12
hidden: displayed_only=False:
Sku Stock
0 AF-1 5
1 AF-2 0
2 AF-3 12
six ways to hide a row, flavor='lxml', displayed_only=True:
Row How hidden
0 1 visible
1 3 inline, upper case
2 4 class + stylesheet
3 5 hidden attribute
4 6 visibility:hidden

A nested table is returned twice: once as its own DataFrame, and once flattened into the outer table's cell, where its text is concatenated and its rows spill into extra outer rows, because both parsers collect every descendant <tr> of the outer table. Read the inner table by its own id and ignore the outer one, or walk the outer table with BeautifulSoup:

nested: attrs id=outer:
Warehouse Bins
0 Berlin BinCount B140 B215
1 Bin Count
2 B1 40
3 B2 15
4 Lyon none
nested: attrs id=inner:
Bin Count
0 B1 40
1 B2 15
nested: match='Bin' returns both: [(5, 2), (2, 2)]

Which parser is running​

flavor picks the HTML parser: "lxml" (the default, tried first) or "bs4" / "html5lib" (synonyms; BeautifulSoup with the html5lib tree builder). On this packet's broken table (unclosed <td>s, a missing </tr>) all three produced the same DataFrame; the IO guide's "HTML Table Parsing Gotchas" section warns that lxml makes no guarantees on invalid markup and that pandas falls back to bs4 when lxml fails to parse (in the code, when lxml raises ValueError; an ImportError does not fall back). The missing-library errors are explicit:

--- flavor with the library missing (subprocess, module blocked)
lxml blocked, flavor='lxml': exit 1 | ImportError: `Import lxml` failed. Use pip or conda to install the lxml package.
bs4 blocked, flavor='bs4': exit 1 | ImportError: `Import beautifulsoup4` failed. Use pip or conda to install the beautifulsoup4 package.
html5lib blocked, flavor='bs4': exit 1 | ImportError: `Import html5lib` failed. Use pip or conda to install the html5lib package.

On this fixture the only difference was speed. On a generated table of 20,000 rows and 6 columns (1.9 MB of HTML):

--- generated table: 20000 rows x 6 cols, 1.9 MB of HTML, median of 5 runs
flavor='lxml' median 1.27 s min 1.20 s shape (20000, 6) dtypes {'id': 'int64', 'name': 'str', 'amount': 'int64', 'bucket': 'int64', 'tag': 'str', 'flag': 'bool'}
flavor='bs4' median 17.19 s min 15.83 s shape (20000, 6) dtypes {'id': 'int64', 'name': 'str', 'amount': 'int64', 'bucket': 'int64', 'tag': 'str', 'flag': 'bool'}

lxml was about 13 times faster on this machine for the same result. Install lxml and leave flavor alone unless lxml misreads a specific page.

Fetching the page: 403s, headers, requests​

Given a URL, pandas fetches it with urllib and its default user agent. Wikipedia refused that on the test date (both calls below pass match="Population", so the shape is the first table containing that word, not tables[0]):

pd.read_html(WIKI) with pandas' default user agent: urllib.error.HTTPError: HTTP Error 403: Forbidden
pd.read_html(WIKI, storage_options=UA): (241, 6)

storage_options (pandas 2.1 and later) forwards its keys as request headers for HTTP URLs, so a User-Agent fixes that case. For anything beyond static headers (a session, retries, timeouts, a proxy), fetch with requests and pass the body through StringIO. The same applies to the BeautifulSoup hand-off from older tutorials, where str(table) used to be passed straight in:

r = requests.get(URL, headers=UA, timeout=30)
soup = BeautifulSoup(r.text, "lxml")
target = soup.find("table", id="prices")
requests.get: 200 text/html; charset=utf-8
pd.read_html(StringIO(r.text)): 12
pd.read_html(r.text) (old idiom): FileNotFoundError: [Errno 2] No such file or directory: <!DOCTYPE html>
pd.read_html(str(target)) (old idiom): FileNotFoundError: [Errno 2] No such file or directory: <table class="data" id="prices">
pd.read_html(StringIO(str(target))):
Product Price Units sold Share
0 Ant Farm Deluxe $1,234.50 12345 45.5%
1 Ant Farm Mini $99.00 1020 3.8%
2 Magnifier $12.25 987654 50.7%
rows by hand for comparison: [['Product', 'Price', 'Units sold', 'Share'], ['Ant Farm Deluxe', '$1,234.50', '12,345', '45.5%']]

If read_html cannot make sense of a table (merged cells that mean something, data in attributes, tables built from <div>s), walk it by hand; Web Scraping HTML Tables with Python does that with BeautifulSoup.

Idioms that no longer run​

The 2024 version of this page carried the first four of these (raw strings, df.append, chunksize=, usecols=); positional match is a fifth that other tutorials use. On pandas 3.0.5:

df.append({...}, ignore_index=True) (removed in pandas 2.0): AttributeError: 'DataFrame' object has no attribute 'append'
pd.concat([df, pd.DataFrame([{...}])]):
Name Age City
0 Ada 36 London
1 Grace 45 Arlington
2 Linus 28 Helsinki
3 Ken 60 Murray Hill
pd.read_html(PAGE, chunksize=1000): TypeError: read_html() got an unexpected keyword argument 'chunksize'
pd.read_html(PAGE, usecols=[0, 1]): TypeError: read_html() got an unexpected keyword argument 'usecols'
table('simple')[['Name', 'Age']] instead:
Name Age
0 Ada 36
1 Grace 45
2 Linus 28
pd.read_html(PAGE, 'Ant Farm') (match positionally): TypeError: read_html() takes 1 positional argument but 2 were given
table('simple', header=5) (row beyond the table): ValueError: Passed header=[5], len of 1, but only 4 lines in file
  • DataFrame.append was removed in pandas 2.0; use pd.concat.
  • chunksize and usecols are read_csv parameters; read_html has neither. Select columns after reading; for a table too large to hold, the problem is the page, not the parser.
  • match is keyword-only, as is every parameter after io, since pandas 2.0.
  • Raw HTML strings: StringIO, as above.

Tables that are not in the HTML​

read_html only sees <table> elements in the bytes it is given. If the page builds its table with JavaScript, the response has either no <table> (the "No tables found" error above) or a header-only skeleton (an empty DataFrame, as shown above), and the HTML has to be rendered first. With the ScrapingAnt API, a request with JavaScript rendering through a datacenter proxy costs 10 API credits; setting browser=false performs the request without a headless browser (no JavaScript rendering); the docs describe it for static content, and such a request costs 1 API credit through a datacenter proxy. The rendered body then goes through StringIO like any other response. The packet's last script is written to do that (KEY is SCRAPINGANT_API_KEY from the environment and TARGET the fixture URL); it runs only when the key is set, and the recorded run is the skipped case:

r = requests.get("https://api.scrapingant.com/v2/general", params={"url": TARGET, "browser": "true"}, headers={"x-api-key": KEY}, timeout=90)
print("status:", r.status_code, "Ant-credits-cost:", r.headers.get("Ant-credits-cost"))
tables = pd.read_html(StringIO(r.text), attrs={"id": "prices"})
$ python 10_scrapingant.py
skipped: SCRAPINGANT_API_KEY is not set
exit=0

When the table is in the HTML the server sends, which is every case on this page, no rendering service is needed. Request and response format: https://docs.scrapingant.com/request-response-format. To get the result into a spreadsheet, df.to_excel() or df.to_csv(); Scrape Data From Websites To Excel covers that end.

Limitations​

  • Everything here was measured on pandas 3.0.5; on pandas 2.1 to 2.3 the raw-string form still works with a FutureWarning (documented, not run), and storage_options needs 2.1 or later.
  • The Wikipedia 403 is what Wikipedia returned to pandas' default user agent on 2026-09-17; it is their policy, not a pandas behaviour, and can change.
  • The parser timing is one machine; the ratio between lxml and bs4 is the claim.
  • 10_scrapingant.py was skipped in the recorded run because no API key was set.

Examples tested on 2026-09-17 with Python 3.12.11, pandas 3.0.5, numpy 2.5.3, lxml 6.1.3, beautifulsoup4 4.15.0, html5lib 1.1, requests 2.34.2. Code: https://github.com/ScrapingAnt/scrapingant-examples/tree/main/examples/pandas-read-html-table.

This article was drafted with AI assistance from a tested evidence packet and reviewed by the named author, who is responsible for the code, measurements and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster