Cheerio Web Scraping in Node.js: A Step-by-Step Tutorial

Published:

14 minute read

Acar Diveroli
Written by: Acar Diveroli
A node scrape-books.mjs command above four worker lanes with one retry, feeding a JSON book record marked 32 BOOKS.

Your team already runs Node.js, and someone asks for the titles, prices and stock counts of every book in one category of a catalogue. The page source in the browser shows the data sitting in plain <article> tags, so you do not need a browser to read it. You need a way to download the HTML, pick the right elements, follow the "next" link, and stop the script from hammering the site or dying on the first timeout. The general route from a page to a file is in How to Extract Data From a Website; this tutorial does it in JavaScript with Cheerio.

We cover what Cheerio is and what it is not, loading HTML with load and fromURL, selectors and lists, the newer extract method, pagination, a concurrency limit, retries with backoff, writing JSON, and sending the requests through a proxy with undici's ProxyAgent. The last part explains when Cheerio is the wrong tool. Every sample ran on 28 September 2026 with cheerio 1.2.0, undici 8.11.2 and Node.js 24.11.1 against books.toscrape.com, a sandbox built for scraping practice.

What is Cheerio?

Cheerio is an HTML and XML parser for Node.js with an API modelled on jQuery. You give it markup, it builds a document tree, and you query that tree with $("css selector"). It is fast because it skips everything a browser does after parsing: no layout, no CSS, no images, no script execution.

The current release is 1.2.0 (cheerio on npm), and it needs Node.js 20.18.1 or later. Version 1.0, released in August 2024, ended a release-candidate phase that began in 2017. The package has no default export, so you write import * as cheerio from "cheerio". Tutorials that call require("cheerio").default were written for older versions.

Cheerio sees only the HTML the server sent. If the product list is filled in later by JavaScript, the data is not in that HTML and no selector will find it. Static vs Dynamic Pages shows how to check which kind of page you have before you write any code.

How does a Cheerio scraper work?

A scraper built on Cheerio repeats the same five steps for each page:

  1. Download. An HTTP client (fetch here) requests the URL and receives the HTML as text.
  2. Parse. cheerio.load(html) builds the tree and returns a $ function bound to that document.
  3. Select. $("article.product_pod") returns every matching element; .find(), .text() and .attr() read inside them.
  4. Follow. The scraper reads the next URL from the page (a pagination link or a detail link) and resolves it against the current URL.
  5. Store. Rows are collected in memory and written to a file or database at the end.

Steps 2 and 3 never touch the network. That split matters for debugging: if a selector returns nothing, save the HTML to a file and test the selector against it, without sending another request.

Cheerio vs jsdom vs Playwright

The three tools most often compared for scraping in Node.js do different jobs:

ToolWhat it doesRuns page JavaScriptCost per pageGood fit
CheerioParses HTML, jQuery-style queriesNoLowest: parse onlyServer-rendered HTML, large page counts
jsdomBuilds a browser-like DOM in Node.jsOptional, limitedHigher than CheerioCode that expects document and DOM APIs
PlaywrightDrives a real Chromium, Firefox or WebKitYesHighest: full browserPages that build content with JavaScript, clicks, logins

A common setup uses both ends: Playwright for the few pages that need a browser, Cheerio for everything else. Language choice is a separate question, covered in Web Scraping: JavaScript or Python?.

Installing Cheerio and loading your first page

Create a project and install the package. Adding "type": "module" lets you use import and top-level await:

bash
mkdir book-scraper && cd book-scraper
npm init -y
npm pkg set type=module
npm install cheerio

Node.js 18 and later ship fetch, so the first script needs nothing else:

js
import * as cheerio from "cheerio";

const url = "https://books.toscrape.com/";
const response = await fetch(url, {
  headers: { "user-agent": "book-research/1.0 (+mailto:you@example.com)" },
});
if (!response.ok) throw new Error(`HTTP ${response.status} for ${url}`);

const $ = cheerio.load(await response.text());

console.log($("title").text().trim());
console.log($("article.product_pod").length, "books on this page");

$("article.product_pod").slice(0, 3).each((i, el) => {
  const card = $(el);
  const title = card.find("h3 a").attr("title");
  const price = card.find(".price_color").text();
  console.log(i + 1, title, price);
});
text
All products | Books to Scrape - Sandbox
20 books on this page
1 A Light in the Attic £51.77
2 Tipping the Velvet £53.74
3 Soumission £50.10

The title comes from the title attribute of the link, not its text: on this site the visible link text is cut short ("In a Dark, Dark ...") while the attribute holds the full title. Check both in the page source before you pick one. The user-agent header names your script and gives the site owner a way to reach you.

Loading methods

Cheerio 1.x has five ways to load a document (Cheerio loading docs):

MethodInputWhen to use it
load(html)A stringYou downloaded the page yourself (the usual case)
loadBuffer(buffer)Raw bytesThe encoding is unknown; Cheerio sniffs it
stringStream(options, cb)Decoded text streamLarge files with a known encoding
decodeStream(options, cb)Raw byte streamLarge files with an unknown encoding
fromURL(url, options)A URLQuick scripts; Cheerio downloads the page itself

fromURL is convenient, but it opens its own undici client for the page's origin. In our test it ignored a dispatcher passed in requestOptions and connected directly, even when that dispatcher pointed at a proxy that rejected every request. For anything that needs a proxy, retries or timeouts, download with fetch and use load.

Selecting elements and reading values

Most scraping code uses a small part of the API:

  • $(selector) selects from the whole document; el.find(selector) searches inside one element.
  • .text() returns the combined text of the selection; .attr("href") returns one attribute of the first element.
  • .each((i, el) => …) loops; .map((i, el) => value).get() turns a selection into a plain array.
  • .first(), .eq(n) and .slice(a, b) narrow a selection.

A selector that matches nothing does not throw. .text() returns an empty string and .attr() returns undefined, so a changed class name produces empty fields rather than an error. Validate the rows you collect (more in the mistakes list below). Selector syntax and why Cheerio has no XPath are in CSS Selector vs XPath.

The extract method

Cheerio 1.0 added $.extract(), which describes the whole record as one object (Cheerio extract docs). A string yields the text of the first match, square brackets collect every match, and { selector, value } reads a property or runs a function:

js
import * as cheerio from "cheerio";

const $ = await cheerio.fromURL("https://books.toscrape.com/");

const data = $.extract({
  heading: "h1",
  books: [
    {
      selector: "article.product_pod",
      value: {
        title: { selector: "h3 a", value: "title" },
        price: ".price_color",
        link: { selector: "h3 a", value: "href" },
        rating: {
          selector: "p.star-rating",
          value: (el) => $(el).attr("class").replace("star-rating", "").trim(),
        },
      },
    },
  ],
});

console.log(data.heading, data.books.length);
console.log(data.books[0]);
text
All products 20
{
  title: 'A Light in the Attic',
  price: '£51.77',
  link: 'catalogue/a-light-in-the-attic_1000/index.html',
  rating: 'Three'
}

Selectors inside value run relative to each article, which keeps the fields of one book together. The link stays relative, so resolve it with new URL(link, pageUrl) before you request it.

A complete scraper: pagination, concurrency, retries and JSON

The script below collects every book in the Mystery category. It walks the listing pages by following the "next" link, opens each book's detail page with at most four requests in flight, retries network errors, timeouts, 429 and 5xx responses with exponential backoff, and writes books.json. It uses undici's fetch so that the optional proxy in the next section works without changes. Install it with npm install cheerio undici (undici 8 needs Node.js 22.19 or later).

js
import * as cheerio from "cheerio";
import { fetch, ProxyAgent } from "undici";
import { writeFile } from "node:fs/promises";

const START_URL =
  "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html";
const CONCURRENCY = 4;   // detail pages fetched at the same time
const MAX_RETRIES = 3;   // extra attempts after the first one
const HEADERS = { "user-agent": "book-research/1.0 (+mailto:you@example.com)" };

// Optional proxy: PROXY_URL=http://user:pass@pr.proxynet.io:8000
const dispatcher = process.env.PROXY_URL
  ? new ProxyAgent(process.env.PROXY_URL)
  : undefined;

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
const backoff = (attempt) => 1000 * 2 ** (attempt - 1) + Math.random() * 250;

class HttpError extends Error {
  constructor(status, url) {
    super(`HTTP ${status} for ${url}`);
    this.status = status;
  }
}

async function fetchHtml(url) {
  for (let attempt = 1; ; attempt++) {
    let wait;
    try {
      const res = await fetch(url, {
        headers: HEADERS,
        dispatcher,
        signal: AbortSignal.timeout(15_000),
      });
      if (res.ok) return await res.text();
      const retryable = res.status === 429 || res.status >= 500;
      if (!retryable || attempt > MAX_RETRIES) throw new HttpError(res.status, url);
      const retryAfter = Number(res.headers.get("retry-after"));
      wait = retryAfter > 0 ? retryAfter * 1000 : backoff(attempt);
    } catch (err) {
      if (err instanceof HttpError || attempt > MAX_RETRIES) throw err;
      wait = backoff(attempt); // network error or timeout
    }
    console.warn(`retry ${attempt}/${MAX_RETRIES} in ${Math.round(wait)} ms: ${url}`);
    await sleep(wait);
  }
}

// Run fn over items with at most `limit` calls in flight.
async function mapLimit(items, limit, fn) {
  const results = new Array(items.length);
  let next = 0;
  async function worker() {
    while (next < items.length) {
      const i = next++;
      results[i] = await fn(items[i], i);
    }
  }
  await Promise.all(Array.from({ length: Math.min(limit, items.length) }, worker));
  return results;
}

function parseListPage(html, pageUrl) {
  const $ = cheerio.load(html);
  const books = $("article.product_pod")
    .map((_, el) => {
      const card = $(el);
      const link = card.find("h3 a");
      return {
        title: link.attr("title"),
        price: Number(card.find(".price_color").text().replace(/[^0-9.]/g, "")),
        rating: card.find("p.star-rating").attr("class").split(" ").pop(),
        url: new URL(link.attr("href"), pageUrl).href,
      };
    })
    .get();
  const nextHref = $("li.next a").attr("href");
  return { books, nextUrl: nextHref ? new URL(nextHref, pageUrl).href : null };
}

function parseDetailPage(html) {
  const $ = cheerio.load(html);
  const info = {};
  $("table.table-striped tr").each((_, row) => {
    info[$(row).find("th").text().trim()] = $(row).find("td").text().trim();
  });
  const stock = info["Availability"]?.match(/\((\d+) available\)/);
  return {
    upc: info["UPC"],
    inStock: stock ? Number(stock[1]) : 0,
    description: $("#product_description + p").text().trim(),
  };
}

// 1. Walk the listing pages by following the "next" link.
const listed = [];
for (let url = START_URL; url; ) {
  const { books, nextUrl } = parseListPage(await fetchHtml(url), url);
  listed.push(...books);
  console.log(`${url} -> ${books.length} books`);
  url = nextUrl;
}

// 2. Open every detail page, four at a time.
const books = await mapLimit(listed, CONCURRENCY, async (book) => {
  try {
    return { ...book, ...parseDetailPage(await fetchHtml(book.url)) };
  } catch (err) {
    console.error(`skipped ${book.url}: ${err.message}`);
    return { ...book, error: err.message };
  }
});

// 3. Save the result.
await writeFile(
  "books.json",
  JSON.stringify({ scrapedAt: new Date().toISOString(), count: books.length, books }, null, 2),
);
console.log(`saved ${books.length} books to books.json`);
text
https://books.toscrape.com/catalogue/category/books/mystery_3/index.html -> 20 books
https://books.toscrape.com/catalogue/category/books/mystery_3/page-2.html -> 12 books
saved 32 books to books.json

One record from books.json (description shortened):

json
{
  "title": "Sharp Objects",
  "price": 47.82,
  "rating": "Four",
  "url": "https://books.toscrape.com/catalogue/sharp-objects_997/index.html",
  "upc": "e00eb4fd7b871a48",
  "inStock": 20,
  "description": "…"
}

What each part does:

  • Pagination. The loop stops when the page has no li.next a. The link on page 1 is page-2.html, relative to the category folder, which is why every URL goes through new URL(href, pageUrl). Other patterns (page numbers in the query string, cursors, "load more" APIs) are covered in Pagination in Web Scraping.
  • Concurrency limit. mapLimit starts four workers that take the next item from a shared counter. Promise.all over all 32 URLs would send 32 requests at once; with 1,000 URLs it would look like a burst to the server. Four is a polite starting point for a small site.
  • Retries. Only errors that can pass on their own are retried: network failures, the 15-second timeout, 429 and 5xx. A 404 fails at once. A numeric Retry-After header wins over the computed delay; the backoff doubles from about one second and adds random jitter so parallel workers do not retry in step. Why 429 happens and how to read the header is in HTTP 429 Too Many Requests.
  • Partial failure. A detail page that still fails after three retries becomes a row with an error field instead of stopping the run. You can re-run only those rows later.
  • JSON. The file carries scrapedAt and count, which helps when you compare runs. For CSV, JSON Lines or SQLite with upserts, see How to Save Scraped Data to CSV, JSON and SQLite.

Using a proxy with Cheerio (undici ProxyAgent)

Cheerio never opens a connection in this script, so the proxy belongs to the HTTP client. With undici, you create a ProxyAgent and pass it to fetch as dispatcher. The full script above already does this when PROXY_URL is set:

bash
PROXY_URL=http://user:pass@pr.proxynet.io:8000 node scrape-books.mjs

undici builds the Proxy-Authorization header from the user name and password in the URL, and URL-decodes them first, so special characters in a password must be percent-encoded (undici ProxyAgent docs). For HTTPS targets the agent opens a CONNECT tunnel, and TLS to the site runs inside it.

We ran the script through a small local proxy that required user:pass and logged every tunnel. All 34 requests (two listing pages, 32 detail pages) arrived through one CONNECT books.toscrape.com:443: the agent kept the tunnel open and reused it. With a wrong password the proxy answered 407, and undici reported Proxy response (407) !== 200 when HTTP Tunneling. The script retried that three times before giving up; a wrong password never fixes itself, so check the credentials rather than raising the retry count.

If you would rather not add undici as a dependency, Node.js 24.5 and 22.21 added built-in proxy support that reads HTTP_PROXY, HTTPS_PROXY and NO_PROXY when you set NODE_USE_ENV_PROXY=1 (Node.js built-in proxy support). The documentation marks it as active development. In our test on Node.js 24.11.1, plain global fetch went through the local proxy with this setting:

bash
NODE_USE_ENV_PROXY=1 HTTPS_PROXY=http://user:pass@pr.proxynet.io:8000 node first-page.mjs

Axios and node-fetch use agents instead of dispatchers; Using a Proxy in Node.js covers both. For rotating exit IPs per request or sticky sessions that keep one IP for a while, Residential Proxy and Rotating Proxy accept the same user:pass@host:port URL.

When Cheerio is not enough

Cheerio cannot click, scroll or wait for a request the page makes after loading. Signs that you need a browser:

  • The page source (Ctrl+U) lacks the data that the rendered page shows.
  • The HTML contains an empty container such as <div id="root"></div> and a large script bundle.
  • The data appears only after a login form, a cookie banner or an "infinite scroll".

Before starting a browser, open the Network tab in developer tools. Many "dynamic" pages load their data from a JSON endpoint, and requesting that endpoint with fetch is lighter than rendering the page. If you do need a browser, Playwright can render the page and hand the final HTML to Cheerio with cheerio.load(await page.content()), so your parsing code stays the same.

Where Cheerio scrapers are used

  • Price tracking: reading prices from server-rendered product pages on a schedule (price monitoring).
  • Catalogue and market data: collecting product ranges and stock levels across shops (market research).
  • Search visibility: checking titles, meta tags and headings on your own pages (SEO proxy).
  • Data pipelines: feeding parsed rows into a larger crawler or ETL job (data scraping, web crawler).
  • Parsing saved HTML: turning archived pages into structured records; the parsing side is explained in What Is Data Parsing?.

Common mistakes and how to diagnose them

  • Empty strings everywhere. The selector matched nothing, or the data is added by JavaScript. Save the HTML with writeFile("page.html", html) and search it for a value you can see in the browser.
  • require(...).default is not a function or does not provide an export named 'default'. Old import style. Use import * as cheerio from "cheerio".
  • TypeError: fetch failed with invalid onRequestStart method. You passed a ProxyAgent from the npm undici package to Node.js's global fetch. Node 24.11.1 bundles undici 7.16.0, and the two versions do not share a dispatcher interface. Import fetch and ProxyAgent from the same package.
  • A proxy that "does nothing". You passed the agent to cheerio.fromURL, which uses its own client. Download with fetch and call cheerio.load.
  • Relative links fail. fetch("catalogue/…") throws Failed to parse URL. Resolve with new URL(href, pageUrl).
  • Too many requests at once. Promise.all(urls.map(fetch)) sends everything in parallel and invites 429 responses. Use a limit such as mapLimit.
  • Silent data drift. The site renames a class and prices become NaN. Check each run: count rows, count NaN prices, and stop if the numbers drop sharply.

Before scaling up, read the site's robots.txt and terms, prefer an official API when there is one, and keep request rates modest. robots.txt explained and Is Web Scraping Legal? cover the rules; Web Scraping Without Getting Blocked covers polite crawling.

Decision guide

NeedRecommendation
Data is in the page sourcefetch + cheerio.load
One-off script, no proxycheerio.fromURL
Many records with the same shape$.extract with an array descriptor
Hundreds of pagesA concurrency limit of 2-5 plus retries with backoff
Requests through a proxyundici fetch + ProxyAgent, or NODE_USE_ENV_PROXY=1 on Node.js 24.5+
Data appears only after JavaScript runsFind the JSON endpoint first, otherwise Playwright + Cheerio
Code expects a full DOM (document, events)jsdom

Frequently asked questions

Is Cheerio still maintained in 2026?

Yes. The npm registry lists version 1.2.0, published in January 2026, as the latest release, and the documentation site covers the 1.x API including extract and fromURL.

Does Cheerio run JavaScript?

No. It parses the HTML string you give it and nothing more. Scripts in the page are treated as text. For pages that build their content in the browser, use Playwright or find the data endpoint the page calls.

Do I need Axios with Cheerio?

No. Node.js 18 and later include fetch, and it covers what most scrapers need. Axios is a matter of taste; if you use it, pass the response body (response.data) to cheerio.load.

How do I scrape multiple pages with Cheerio?

Read the next-page link from each page, resolve it against the current URL, and loop until the link is missing, as in the full script above. When the page count is known, you can also build the URL list up front and run it through a concurrency limit.

How do I use a proxy with Cheerio?

Configure the proxy on the HTTP client, not on Cheerio. With undici: new ProxyAgent("http://user:pass@pr.proxynet.io:8000"), passed as dispatcher to undici's fetch. On Node.js 24.5 or later you can instead set NODE_USE_ENV_PROXY=1 and HTTPS_PROXY.

Is Cheerio faster than Puppeteer or Playwright?

For pages whose data is in the HTML, yes, because it only parses text while a browser also downloads assets, runs scripts and lays out the page. We did not benchmark the difference, and it depends on the page, so measure on your own targets if the numbers matter.

Summary

Cheerio turns downloaded HTML into a tree you can query with CSS selectors, and in version 1.x it adds fromURL and extract. A dependable scraper keeps the network work outside Cheerio: fetch with a timeout, retries for 429, 5xx and network errors, a small concurrency limit, and a JSON file with a timestamp. When requests need to leave from another IP or country, pass an undici ProxyAgent as the dispatcher and point it at a Proxynet proxy.

Ask ChatGPTAsk Claude