---
title: "What Is Data Parsing? Parser Types, Errors and Examples"
description: "Data parsing turns raw input such as HTML, JSON, CSV or logs into structured records your code can use. How parsers work, the main types and the usual errors."
url: https://proxynet.io/blog/what-is-data-parsing
date: 2026-09-28
author: "Acar Diveroli"
category: "Web Scraping, Proxies"
lang: en
---

# What Is Data Parsing? Parser Types, Errors and Examples

A scraper downloads a product page and saves it to disk. The file is 180 KB of HTML, and somewhere inside it are the four things you actually wanted: the product name, the price, the stock status and the SKU. Until a program finds those values, checks them and writes them into named fields, the download is just text. That step, from raw text to fields, is data parsing.

This post explains what parsing means, how a parser works (two core stages plus the steps around them), and which kind of parser fits HTML, JSON, CSV, logs and custom formats. It separates parsing from scraping and extraction, shows a short tested Python example with BeautifulSoup and the `json` module, covers the errors that break parsers in real projects (malformed HTML, invalid JSON, schema drift), and ends with a build-or-buy guide.

> **Note: Short answer**
>
> Data parsing is the process of reading raw input, such as an HTML page, a JSON response, a CSV file or a log line, and turning it into a structure a program can work with: objects, rows or key-value pairs. A parser first splits the input into tokens, then arranges the tokens according to the rules of the format. In web scraping, the scraper fetches the page and the parser turns that page into clean records, which are then validated and stored.

## What is data parsing?

Parsing means analyzing a piece of input according to a set of rules, and producing a structure that represents what the input says. The input is usually a string or a stream of bytes. The output is something with names and types: a tree of HTML elements, a Python dictionary, a list of rows, a record with a `price` field that holds a number instead of the text `"$1,249.00"`.

Your browser parses HTML into a page, your API client parses JSON into objects, your spreadsheet parses CSV into cells. You notice parsing only when the input does not match what the parser expected, and the program either stops with an error or, worse, keeps going with wrong values.

In data work the word is used in two layers:

- **Format parsing.** Turning bytes into the format's own structure: HTML into a DOM tree, JSON into objects, CSV into rows. Libraries do this.
- **Data parsing in the narrow sense.** Taking that structure and pulling out the values you need, with the right types: the price as a decimal, the date as a date, the stock status as true or false. You usually write this part yourself.

## How does a parser work?

Almost every parser, from a JSON library to a browser's HTML engine, follows the same two core stages, lexing and parsing, with a read step before them and value selection and validation after.

1. **Read the input as text.** Bytes are decoded into characters with an encoding, normally UTF-8. A wrong guess here produces garbled characters before parsing even starts; [Python Unicode Encoding Errors](/blog/python-unicode-encoding-errors) covers that failure.
1. **Tokenize (lexical analysis).** The parser cuts the character stream into tokens, the smallest meaningful pieces: a start tag, an attribute, a string, a number, a comma, a closing brace.
1. **Build the structure (syntactic analysis).** The tokens are arranged according to the grammar of the format. For HTML the result is a tree of nested elements; for JSON, nested objects and arrays; for CSV, rows of fields.
1. **Select the values.** Your code walks the structure and picks the fields it needs, for example with a CSS selector or XPath on an HTML tree. [CSS Selector vs XPath](/blog/css-selector-vs-xpath) compares the two ways of pointing at an element.
1. **Convert and validate.** Text becomes typed data: `"$34.50"` becomes `34.50`, `"In stock"` becomes `true`. Records missing a required field are flagged instead of silently saved.

The HTML standard describes a browser's parser with exactly these stages. The [WHATWG parsing chapter](https://html.spec.whatwg.org/multipage/parsing.html) splits it into tokenization and tree construction, and it also defines how a parser must recover from errors, which is why browsers show broken pages instead of refusing them.

## Parsing vs scraping vs extraction

The three words are often used as synonyms, but they name different steps of one pipeline.

| Term | What it does | Input | Output |
|---|---|---|---|
| Crawling | Finds pages by following links | A start URL | A list of URLs |
| Scraping | Downloads the content of those pages | URLs | Raw HTML, JSON or files |
| Parsing | Turns raw content into structure | Raw text | Trees, objects, rows |
| Extraction | Picks the needed values out of the structure | A tree or object | Named fields |
| Validation and cleaning | Checks types, removes duplicates, fixes formats | Fields | Records you can trust |
| Storage | Writes the records somewhere | Records | CSV, JSON Lines, a database |

In everyday speech "parsing" covers parsing, extraction and part of validation together, as in the rest of this post. Crawling and scraping are the fetching side; [Web Scraping vs Web Crawling](/blog/web-scraping-vs-web-crawling) explains that split. Where the records end up is covered in [How to Save Scraped Data to CSV, JSON and SQLite](/blog/save-scraped-data-csv-json-sqlite), and the whole chain as one repeatable job is the subject of [What Is ETL?](/blog/what-is-etl).

## Types of parsers

Different inputs call for different parsers. Using the wrong kind, for example a regular expression on nested HTML, is the source of many fragile scrapers.

| Parser type | Best for | Examples | Weak spot |
|---|---|---|---|
| HTML / DOM parser | Web pages, including broken markup | BeautifulSoup with `html.parser`, `lxml` or `html5lib`; Cheerio in Node.js | Needs selectors that change when the page layout changes |
| JSON parser | API responses, embedded page data | Python `json`, `JSON.parse` in JavaScript | Strict: one stray comma fails the whole document |
| CSV / delimited parser | Exports, spreadsheets, reports | Python `csv`, pandas `read_csv` | Quoting, delimiters and encodings vary by producer |
| XML parser | Feeds, sitemaps, older enterprise APIs | `lxml`, `xml.etree.ElementTree` | Namespaces make selectors verbose |
| Regular expressions | Small flat patterns inside a known field: prices, dates, IDs | Python `re` | Cannot follow nesting; breaks on layout changes |
| Grammar-based parser | Custom formats, query languages, config files | Parser generators such as ANTLR, Lark | Takes time to write the grammar |
| Model-based parser | Messy text with no fixed layout | Language models that return JSON | Output must be validated; results can vary between runs |

Two rows need a note. **HTML parsers differ:** BeautifulSoup sits on top of a parser you choose, and its [documentation](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) shows the broken snippet `<a></p>` producing three different trees in `html.parser`, `lxml` and `html5lib`. Name the parser explicitly, or your output can change on a machine with a different one installed. **Model-based parsing** is how AI scrapers read pages whose layout keeps changing; it trades exact rules for flexibility, so validation matters even more, as [How AI Web Scrapers Work](/blog/ai-web-scraper-how-it-works-2026) explains.

## A short parsing example in Python

The example parses a small product page with an HTML parser, CSS selectors, one regular expression for the price field, a required-field check and a separate JSON parse for the data embedded in the page. Save this as `page.html`; the second card has an unclosed `<span>` and the third has no price:

```html
<!doctype html>
<html lang="en">
<head>
  <title>Desk lamps</title>
  <script type="application/ld+json">
  {"@context": "https://schema.org", "@type": "ItemList",
   "itemListElement": [
     {"@type": "ListItem", "position": 1, "url": "/p/arc-lamp"},
     {"@type": "ListItem", "position": 2, "url": "/p/clip-lamp"}
   ]}
  </script>
</head>
<body>
  <div class="product" data-sku="L-100">
    <h2 class="name">Arc Desk Lamp</h2>
    <span class="price">$1,249.00</span>
    <span class="stock">In stock</span>
  </div>
  <div class="product" data-sku="L-200">
    <h2 class="name">  Clip Lamp </h2>
    <span class="price">$34.50</span>
    <span class="stock">Out of stock
  </div>
  <div class="product" data-sku="L-300">
    <h2 class="name">Floor Lamp</h2>
    <span class="stock">In stock</span>
  </div>
</body>
</html>
```

Install the two libraries with `pip install beautifulsoup4 lxml`, then save this as `parse_products.py` next to the page:

```python
import json
import re
from decimal import Decimal
from pathlib import Path

from bs4 import BeautifulSoup

PRICE_RE = re.compile(r"[\d.,]+")

def parse_price(text):
    """'$1,249.00' -> Decimal('1249.00'); None if there is no number."""
    match = PRICE_RE.search(text or "")
    if not match:
        return None
    return Decimal(match.group().replace(",", ""))

def parse_products(html):
    soup = BeautifulSoup(html, "lxml")
    rows, problems = [], []

    for card in soup.select("div.product"):
        name = card.select_one(".name")
        price = card.select_one(".price")
        stock = card.select_one(".stock")
        row = {
            "sku": card.get("data-sku"),
            "name": name.get_text(strip=True) if name else None,
            "price": parse_price(price.get_text()) if price else None,
            "in_stock": stock is not None and stock.get_text(strip=True) == "In stock",
        }
        missing = [key for key in ("sku", "name", "price") if row[key] is None]
        if missing:
            problems.append({"sku": row["sku"], "missing": missing})
            continue
        rows.append(row)

    # Structured data embedded in the page: parse it as JSON, not as HTML.
    urls = []
    for tag in soup.select('script[type="application/ld+json"]'):
        try:
            data = json.loads(tag.string or "")
        except json.JSONDecodeError as err:
            problems.append({"json_ld": f"line {err.lineno}, col {err.colno}: {err.msg}"})
            continue
        urls += [item.get("url") for item in data.get("itemListElement", [])]

    return rows, urls, problems

if __name__ == "__main__":
    html = Path("page.html").read_text(encoding="utf-8")
    rows, urls, problems = parse_products(html)
    print(json.dumps(rows, indent=2, default=str))
    print("urls:", urls)
    print("problems:", problems)
```

Running `python parse_products.py` with BeautifulSoup 4.15.0, lxml 6.1.3 and Python 3.13 prints:

```text
[
  {
    "sku": "L-100",
    "name": "Arc Desk Lamp",
    "price": "1249.00",
    "in_stock": true
  },
  {
    "sku": "L-200",
    "name": "Clip Lamp",
    "price": "34.50",
    "in_stock": false
  }
]
urls: ['/p/arc-lamp', '/p/clip-lamp']
problems: [{'sku': 'L-300', 'missing': ['price']}]
```

The unclosed `<span>` broke nothing because `lxml` repaired the tree, the padded name came out clean, and the lamp with no price went to `problems` instead of being saved empty. When we added a trailing comma to the JSON-LD block as a test, the script kept running and reported `line 5, col 64: Illegal trailing comma before end of object`, which points straight at the fault. [JSONDecodeError: Expecting Value](/blog/jsondecodeerror-expecting-value) lists the other messages this parser raises and what causes each one.

The same parser written for Node.js uses Cheerio instead of BeautifulSoup; see [Cheerio Web Scraping](/blog/cheerio-web-scraping). For a full scraper with requests, pagination and saving, start with the [BeautifulSoup tutorial](/blog/beautifulsoup-tutorial).

## Common parsing errors and what causes them

- **Malformed HTML.** Unclosed tags, stray closing tags and tags in the wrong place are normal on the web. Browsers and HTML parsers repair them, but each parser repairs them its own way, so a selector that works with `html5lib` may miss with `html.parser`. Fix the parser choice and test against saved pages.
- **Invalid JSON.** JSON is strict. [RFC 8259](https://www.rfc-editor.org/rfc/rfc8259) defines the grammar, and trailing commas, single quotes and comments are not part of it. A frequent cause is not JSON at all: the server returned an HTML error page or an empty body, and the JSON parser fails on the first character.
- **Duplicate keys.** RFC 8259 says that when names in an object repeat, how the receiver behaves is unpredictable. Python's [json module](https://docs.python.org/3/library/json.html) keeps the last value without raising an error, so `{"price": 10, "price": 12}` quietly becomes `12`.
- **CSV quoting and delimiters.** A comma inside a product name splits one field in two unless it is quoted. [RFC 4180](https://www.rfc-editor.org/rfc/rfc4180) describes the usual quoting convention, but it is only informational and many exports use semicolons or tabs. Use a CSV library, never `line.split(",")`.
- **Locale-specific numbers and dates.** `1.249,00` is over a thousand in Germany and about one in a naive parser; `03/04/2026` is March or April depending on the country. Parse with an explicit format per source.
- **Schema drift.** The site renames a class or moves the price into a new element, and the parser keeps running but returns `None` or the wrong value. Nothing crashes, which makes this the most expensive failure. Required-field checks, like the `problems` list above, turn silent drift into a visible count.
- **Content that is not in the HTML.** Some pages build their content with JavaScript after loading, so the data was never in the downloaded file. [Static vs Dynamic Pages](/blog/static-vs-dynamic-pages) shows how to tell.

## Should you build or buy a parser?

For standard formats you never write the low-level parser; the libraries are mature. The decision is about the extraction layer on top: selectors, conversions and checks per site.

**Build it yourself when** you have a limited number of sources that change rarely. Your own code is transparent, costs nothing per page and is easy to test against saved samples.

**Use a ready-made tool or service when** you track hundreds of sites whose layouts change often, or when the input has no fixed layout at all, such as emails and PDFs. You pay per volume or per seat and accept less control over the output.

**A hybrid is common:** official APIs or embedded JSON where they exist, hand-written parsers for the key sources, a model-based fallback for the long tail, and the same validation step in front of storage for all three.

## Where parsing is used

- **Data collection at scale.** Every scraping pipeline parses between download and database; see [data scraping](/data-scraping).
- **Crawlers and site audits.** A crawler parses each page for links and metadata; see [web crawler](/web-crawler).
- **Price tracking.** Prices parsed into numbers can be compared over time, as in [competitor price tracking](/blog/competitor-price-tracking).
- **Change monitoring.** Comparing parsed fields instead of raw HTML avoids false alarms from ads and timestamps; see [Website Change Monitoring](/blog/website-change-monitoring).
- **Analytics.** Typed records are where pattern finding starts; see [What Is Data Mining?](/blog/what-is-data-mining).
- **Logs and security.** Log lines are parsed into timestamp, IP address, status and path before anyone can count errors.

## Common mistakes

- **Parsing the live site while you develop.** Save a few real pages to disk and write the parser against them: fewer requests to the site, repeatable tests.
- **Using regex on whole HTML documents.** Parse the tree first, then use regex only inside a single field's text.
- **Storing prices as text or floats.** Convert early to `Decimal` or integer cents and keep the raw string for debugging.
- **Swallowing every exception.** A bare `try/except: pass` hides schema drift. Collect failures with the record ID and check the count after each run.
- **Ignoring the embedded JSON.** Many product pages carry JSON-LD that holds the same values in cleaner form than the visible HTML.
- **Blaming the parser for blocked requests.** When the "HTML" is a block page or a 429 error, no selector will match. Check the status code before parsing, and slow down if the site asks you to.

## Decision guide

| Your input | Recommended approach |
|---|---|
| An API that returns JSON | Use the API; parse with the standard JSON library and validate required fields |
| Static HTML pages from a few sites | BeautifulSoup with `lxml`, or Cheerio in Node.js, plus CSS selectors |
| Pages that render content with JavaScript | Look for embedded JSON or the site's API first; use a headless browser only if needed |
| CSV or spreadsheet exports | A CSV library or pandas, with the delimiter and encoding set explicitly |
| Server logs in a fixed layout | Regex per line or a log parser, then convert fields to types |
| A custom text format you control | A grammar-based parser (ANTLR, Lark) |
| Emails, PDFs, free text with no layout | A model-based extractor with strict output validation |
| Hundreds of sites with changing layouts | A commercial parsing service or a hybrid setup |

## Frequently asked questions

### What is data parsing in simple terms?

It is reading raw text and turning it into labeled pieces a program can use. A web page goes in, and a record with a name, a price and a stock status comes out.

### What is the difference between parsing and scraping?

Scraping downloads content from a website; parsing turns that content into structured data. A scraper without a parser leaves you with raw HTML files.

### Which Python library is used for parsing HTML?

BeautifulSoup is the most widely used and works on top of `html.parser`, `lxml` or `html5lib`. `lxml` can also be used directly and supports XPath. For JSON, the built-in `json` module is enough.

### Can I parse HTML with regular expressions?

Only for very small, flat pieces of text. HTML is nested and often broken, and a regex cannot follow nesting. Parse the page with an HTML parser, then use regex inside a single field if you need to.

### Why does my parser suddenly return empty values?

Usually the site changed its markup and your selectors no longer match. Other causes: an error page instead of the content, or content loaded later by JavaScript.

### Do I need a proxy to parse data?

Parsing itself runs on your own machine and needs no proxy. Proxies matter on the fetching side, when a scraper collects public pages at volume or needs to see a page as visitors in another country see it. [Residential Proxy](https://proxynet.io/residential-proxy) addresses let you choose the country and city, and [Rotating Proxy](https://proxynet.io/rotating-proxy) spreads requests across addresses; keep to each site's terms and rate limits either way.

## Summary

Data parsing turns raw input into structured, typed records: a tokenizer cuts the text into pieces, a parser arranges them by the format's rules, and your code picks, converts and checks the values it needs. Choose the parser by the input, and keep regex inside single fields. Most of the work in a real project goes into the failures, above all schema drift that no error message announces, so validate every record and count the failures after each run. When the fetching side needs addresses in specific countries, see our [proxy services](/proxy).
