Sentiment Analysis in Python: VADER vs Hugging Face

Published:

14 minute read

Acar Diveroli
Written by: Acar Diveroli
Three tilted review cards with stars and score bars under a fan of blue light beams reaching NEG, NEU and POS badges.

Your team sells three products and has a few thousand reviews for each. The star average barely moves from month to month, yet support says complaints about one product have doubled. Reading every review is not an option, so you want a script that tags each review as positive, neutral or negative and shows where the mood is changing. The first tutorial you find uses VADER, the second a transformer model, and the two give different answers on the same sentence.

This guide runs both methods on the same small set of product reviews. We cover how each one works, the full tested code (a lexicon pass with NLTK, a model pass with the Hugging Face pipeline, and a report that groups results by product and month), how far each method agreed with the star ratings, and the reviews that fooled them. Every script ran on 28 September 2026 with Python 3.13, nltk 3.10.3, pandas 3.0.6, transformers 5.17.0 and the CPU build of torch 2.14.

What is sentiment analysis?

Sentiment analysis, also called opinion mining, is the task of deciding what attitude a text expresses. For product reviews the output is usually one of three labels (positive, neutral, negative) or a star score, often with a confidence score. Aspect-level tools go further and split "great sound, weak battery" into two opinions; that needs a different kind of model and is outside this guide.

Like the methods in What Is Data Mining?, it turns unstructured text into a column you can count and chart. The value comes from the trend, not a single score: the share of negative reviews for one product rising after a supplier change, for example.

How does sentiment analysis work in Python?

There are two families of tools, and they work in very different ways.

Lexicon-based (VADER). VADER, short for Valence Aware Dictionary and sEntiment Reasoner, is a list of about 7,500 English words, emoticons and slang terms that human raters scored from very negative to very positive, plus a set of rules. The rules raise the score for !!! and ALL CAPS, give the words after "but" more weight than those before it, and flip the score when a negation such as "not" or "never" appears within three words before a sentiment word. The vaderSentiment README describes these rules and the thresholds. NLTK ships the same analyzer as nltk.sentiment.vader, documented in the NLTK sentiment how-to.

Model-based (transformers). A transformer such as BERT was first trained on a large amount of text, then fine-tuned on examples labelled with a sentiment. It turns the whole sentence into numbers and reads each word in the light of the others, so "not bad" and "killer" can be judged by context rather than by a fixed value. Hugging Face hosts thousands of such models, and its pipeline function loads one with a single line.

In both cases the scoring runs in five steps:

  1. Load the text. One row per review, with the product, date and, if you have it, the star rating.
  2. Split or tokenize. VADER works on words and punctuation. A transformer uses its own tokenizer, which cuts words into sub-word pieces and stops at a maximum length (512 tokens for BERT-based models).
  3. Score. VADER adds up word values and applies its rules, then squeezes the total into a compound value between -1 and +1. A model returns a probability for each label.
  4. Map to a label. For VADER you pick thresholds. For a model you take the label with the highest probability and keep the score as a confidence.
  5. Aggregate. Count labels per product, week or month, and put the results next to the numbers you already track.

VADER and a transformer model compared

VADER (NLTK)Transformer (Hugging Face)
How it decidesWord list with rated words plus rulesNeural network fine-tuned on labelled text
Install sizenltk plus a 90 KB lexicontransformers, torch (about 500 MB for the CPU build) and a model (670 MB for the one below)
Speed in our test (laptop CPU, short reviews)About 14,000 reviews per secondAbout 90 reviews per second, more on a GPU
LanguagesEnglishDepends on the model; the one below covers six
Context and negationRule-based, three-word windowLearned from examples
Outputneg / neu / pos shares and a compound scoreLabel and probability, or star score
Explaining a resultEasy: look up the wordsHard: needs extra tools
Our 24-review test14 of 24 matched the stars (58 %)19 of 24 matched (79 %)

The last row comes from a tiny, hand-written dataset. It shows the kind of errors each method makes, not how they will do on your data.

Choosing a Hugging Face model for product reviews

Many sentiment models on the Hugging Face hub were trained on tweets or film reviews. Three things matter for product reviews:

  • Training data. A model trained on tweets knows slang; one trained on product reviews knows product vocabulary.
  • Labels. Three labels, five labels or star ratings. Map them to your own labels in one place.
  • Licence. Some popular models carry a non-commercial licence. Read the model card.

We tested three small multilingual models on the same 24 reviews and used nlptown/bert-base-multilingual-uncased-sentiment in the code below. It was fine-tuned on product reviews in English, Dutch, German, French, Spanish and Italian, predicts one to five stars and is MIT licensed. Its model card reports 67 % exact-star accuracy and 95 % within one star for English. Turkish is not among its languages, so a Turkish review gets a guess, not a trained answer.

The other two were lxyuan/distilbert-base-multilingual-cased-sentiments-student (three labels, Apache 2.0), which agreed with our stars on 13 of 24 reviews, and tabularisai/multilingual-sentiment-analysis (five labels, 16 of 24), which is licensed CC BY-NC 4.0 and so not free for commercial use.

The test data: 24 product reviews

We did not use scraped reviews. The file below holds 24 reviews we wrote ourselves for three made-up products, each with a date and a star rating, so both methods can be checked against the rating. We put in the cases that cause trouble in real data: sarcasm, slang, negation, mixed feelings, and one review each in German, Spanish and Turkish. Save it as reviews.csv:

text
review_id,product,date,rating,lang,text
1,kettle,2026-07-02,5,en,"Boils a full litre in under three minutes and the lid never drips. Love it."
2,kettle,2026-07-05,4,en,"Solid kettle, a bit loud but it does the job."
3,kettle,2026-07-11,1,en,"Stopped working after two weeks. Support never answered my emails."
4,kettle,2026-07-19,2,en,"The handle gets hot. Not great, not terrible."
5,kettle,2026-08-03,5,en,"Best purchase this year!!! Fast, quiet and it looks good on the counter :)"
6,kettle,2026-08-14,1,en,"Oh great, another kettle that leaks all over the counter. Just what I needed."
7,kettle,2026-08-21,3,en,"It is a kettle. It boils water."
8,kettle,2026-08-30,4,de,"Kocht schnell und ist leise. Gutes Preis-Leistungs-Verhältnis."
9,headphones,2026-07-03,5,en,"The noise cancelling is superb and the battery lasts all week."
10,headphones,2026-07-09,2,en,"Sound is fine but the ear cushions started peeling after a month."
11,headphones,2026-07-15,4,en,"Comfortable for long flights. The app is clunky though."
12,headphones,2026-07-28,1,en,"Right ear cut out on day three. Returned them."
13,headphones,2026-08-06,5,es,"Suenan de maravilla y la batería dura muchísimo."
14,headphones,2026-08-12,3,en,"Bass is heavy. Some people will like it, I am not sure I do."
15,headphones,2026-08-24,2,en,"I wanted to love these, but the Bluetooth keeps dropping."
16,headphones,2026-08-29,4,en,"Not bad at all for the price."
17,backpack,2026-07-01,5,en,"Survived a rainy week in the mountains and everything inside stayed dry."
18,backpack,2026-07-13,4,en,"Lots of pockets. The laptop sleeve is a little tight for a 16 inch model."
19,backpack,2026-07-22,1,en,"The zipper broke on the first trip. Cheap material."
20,backpack,2026-08-02,3,tr,"Fiyatına göre idare eder ama askılar biraz ince."
21,backpack,2026-08-09,5,en,"Killer design, this bag is sick!"
22,backpack,2026-08-17,2,en,"Looks nice in the photos, feels flimsy in real life."
23,backpack,2026-08-26,4,en,"Carries my camera gear without any back pain."
24,backpack,2026-08-31,1,en,"Terrible stitching. The strap tore off after a week."

Real review data needs a cleaning pass before scoring: duplicates, HTML left in the text, broken characters. How to Clean Scraped Data with Pandas covers that step.

Step 1: score the reviews with VADER

Install the two libraries in a virtual environment:

bash
python -m venv .venv
.venv\Scripts\activate          # macOS/Linux: source .venv/bin/activate
pip install nltk pandas

The script scores each sentence on its own and averages the results. We added that after the first run: scored as one string, review 1 ("…the lid never drips. Love it.") came out at -0.52, because VADER's three-word negation window reached across the full stop and flipped "love". Scored sentence by sentence, it came out at +0.32.

python
import re

import nltk
import pandas as pd
from nltk.sentiment.vader import SentimentIntensityAnalyzer

nltk.download("vader_lexicon", quiet=True)  # 90 KB lexicon, downloaded once

analyzer = SentimentIntensityAnalyzer()
SENTENCE_END = re.compile(r"(?<=[.!?])\s+")


def vader_compound(text: str) -> float:
    """Score each sentence separately and average the compound scores."""
    sentences = [s for s in SENTENCE_END.split(text) if s.strip()]
    scores = [analyzer.polarity_scores(s)["compound"] for s in sentences]
    return round(sum(scores) / len(scores), 4)


def vader_label(compound: float) -> str:
    # thresholds recommended by the VADER authors
    if compound >= 0.05:
        return "positive"
    if compound <= -0.05:
        return "negative"
    return "neutral"


df = pd.read_csv("reviews.csv", parse_dates=["date"])
df["compound"] = df["text"].apply(vader_compound)
df["vader"] = df["compound"].apply(vader_label)

print(df[["review_id", "lang", "compound", "vader"]].to_string(index=False))
df.to_csv("reviews_vader.csv", index=False)

Part of the output:

text
 review_id lang  compound    vader
         1   en    0.3185 positive
         3   en    0.0878 positive
         6   en    0.3125 positive
         8   de    0.0000  neutral
        17   en    0.4588 positive
        21   en   -0.8356 negative

Review 3 is a one-star complaint, but "support" is a positive word in the lexicon and nothing negative outweighs it. Review 6 is sarcasm ("Oh great … Just what I needed") and VADER takes "great" at face value. The German review scores exactly zero, because none of its words are in the English lexicon. Review 21 uses "killer" and "sick" as praise; VADER knows both words only in their negative sense.

Step 2: score the same reviews with a Hugging Face model

The model pass needs transformers and a deep learning backend. If you have no GPU, install the CPU build of torch first; it is much smaller than the default build with CUDA:

bash
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install transformers

The first run downloads the model (about 670 MB) into the Hugging Face cache. Set the HF_HOME environment variable if you want that cache on a different disk. The sentiment-analysis task name is an alias of text-classification; batch_size, truncation and device are described in the Transformers pipelines documentation.

python
import pandas as pd
from transformers import pipeline

MODEL = "nlptown/bert-base-multilingual-uncased-sentiment"

classifier = pipeline(
    "sentiment-analysis",  # alias of "text-classification"
    model=MODEL,
    device="cpu",          # device=0 for the first GPU
)

df = pd.read_csv("reviews_vader.csv", parse_dates=["date"])
results = classifier(
    df["text"].tolist(),
    batch_size=8,
    truncation=True,       # reviews longer than 512 tokens are cut, not rejected
)


def stars_to_label(stars: int) -> str:
    if stars >= 4:
        return "positive"
    if stars == 3:
        return "neutral"
    return "negative"


df["hf_stars"] = [int(r["label"][0]) for r in results]  # "4 stars" -> 4
df["hf_score"] = [round(r["score"], 3) for r in results]
df["hf"] = df["hf_stars"].apply(stars_to_label)

print(df[["review_id", "rating", "hf_stars", "hf_score", "vader", "hf"]].to_string(index=False))
df.to_csv("reviews_scored.csv", index=False)

With the model already in the cache, the script ran in about seven seconds on a laptop CPU, loading included. Part of the output:

text
 review_id  rating  hf_stars  hf_score    vader       hf
         3       1         1     0.736 positive negative
         6       1         1     0.230 positive negative
         8       4         5     0.526  neutral positive
        17       5         1     0.562 positive negative
        20       3         3     0.622  neutral  neutral
        21       5         1     0.816 negative negative

The model caught the complaint in review 3, the sarcasm in review 6 (with a low confidence of 0.23) and the German review. It failed on review 17, the backpack that "survived a rainy week", and read the slang in review 21 as one star with high confidence. Review 20 is Turkish, which the model was not trained on; its three stars matched, but we would not rely on that.

Step 3: aggregate by product and month

The report checks both methods against the star ratings, builds a net sentiment score (share of positive minus share of negative) per product and month, and writes the reviews on which the methods disagree to a file for a person to read.

python
import pandas as pd

df = pd.read_csv("reviews_scored.csv", parse_dates=["date"])


def stars_to_label(rating: int) -> str:
    if rating >= 4:
        return "positive"
    if rating == 3:
        return "neutral"
    return "negative"


df["truth"] = df["rating"].apply(stars_to_label)
for method in ("vader", "hf"):
    hit = (df[method] == df["truth"]).mean()
    print(f"{method}: {hit:.0%} agree with the star rating")
off_by_one = ((df["hf_stars"] - df["rating"]).abs() <= 1).mean()
print(f"model stars within one star of the rating: {off_by_one:.0%}")

# Net sentiment: share of positive minus share of negative, from -1 to +1
NET = {"positive": 1, "neutral": 0, "negative": -1}
df["net"] = df["hf"].map(NET)
df["month"] = df["date"].dt.to_period("M")

by_product = (
    df.groupby("product")
    .agg(reviews=("review_id", "count"),
         net=("net", "mean"),
         negative_share=("hf", lambda s: (s == "negative").mean()),
         avg_rating=("rating", "mean"))
    .round(2)
    .sort_values("net")
)
print(by_product.to_string())

by_month = df.pivot_table(index="month", columns="product", values="net", aggfunc="mean").round(2)
print(by_month.to_string())

# Rows where the two methods disagree go to a person for review
disagree = df[df["vader"] != df["hf"]]
print(f"{len(disagree)} of {len(df)} reviews need a human look")
disagree[["review_id", "rating", "vader", "hf", "text"]].to_csv("review_queue.csv", index=False)
text
vader: 58% agree with the star rating
hf: 79% agree with the star rating
model stars within one star of the rating: 92%
            reviews   net  negative_share  avg_rating
product
backpack          8 -0.25            0.50        3.12
headphones        8  0.12            0.25        3.25
kettle            8  0.12            0.38        3.12
product  backpack  headphones  kettle
month
2026-07     -0.33       -0.25    0.00
2026-08     -0.20        0.50    0.25
11 of 24 reviews need a human look

The backpack and the kettle share an average rating of 3.12, yet half of the backpack reviews read as negative against 38 % for the kettle, a gap the star average hides. With eight reviews per product, one review moves the net score by 0.25, so set a minimum count per group before you chart it. How to store the scored rows for the next run is in How to Save Scraped Data to CSV, JSON and SQLite.

Where do the reviews come from?

The code does not care where the text comes from, but the source decides what you may do with it:

  • Your own shop, app store console or support desk. The cleanest source, and it carries the star rating you need to check the model.
  • Review platforms and marketplaces. Many offer an official API or a seller export. Use it before you consider scraping.
  • Public pages. Follow robots.txt and the terms of service, keep the request rate low and leave out reviewer names and profile links. Is Web Scraping Legal? and Personal Data in Scraped Datasets cover the rules.

When a marketplace shows different reviews or prices per country, a Residential Proxy with country targeting lets you see the page a local shopper sees.

Use cases

  • Product and category research: comparing how buyers talk about your products and a competitor's, as part of market research.
  • E-commerce monitoring: watching review sentiment next to price and stock data across marketplaces, see e-commerce proxies and Competitor Price Tracking.
  • Brand protection: spotting a sudden run of negative reviews that mention counterfeits or unauthorised sellers, see brand protection.
  • Investment research: review and social sentiment is one input in alternative data sets.
  • Support triage: sending one-star and low-confidence reviews to a person first.

Common mistakes and pitfalls

  • Scoring a whole paragraph in one go with VADER. Negation and "but" rules can reach across sentences, as review 1 showed. Score sentences and combine them.
  • Running VADER on other languages. Unknown words score zero, so German and Spanish reviews come out neutral. Use a multilingual model, or group reviews by language and pick a model for each.
  • Trusting sarcasm and slang to either method. Both read "Killer design, this bag is sick!" as a one-star review.
  • Ignoring the domain. "Unpredictable" is good for a film plot and bad for car brakes. Check a sample from your own category.
  • Treating agreement as truth. Review 21 did not appear in the disagreement queue because both methods got it wrong in the same way. Keep a small hand-labelled sample and measure against it.
  • Dropping the confidence score. A label with a score of 0.23, as review 6 got, is barely better than a guess. Send low scores to a person.
  • Forgetting truncation. Without truncation=True a review longer than the model limit raises an error and stops the batch.
  • Charting tiny groups. A net score from five reviews swings wildly. Set a minimum group size.
  • Broken text encoding. Reviews with é in place of é confuse both methods; see Python Unicode Encoding Errors.

Decision guide

NeedRecommendation
A quick first look at English reviewsVADER, scored per sentence
Reviews in several languagesA multilingual model whose card lists your languages
A star score instead of three labelsA model trained on star ratings, such as the nlptown model
No GPU and millions of rowsVADER, or a small model with batching and a sample-based check
Commercial useCheck the model licence; skip non-commercial models
Numbers you will present to managementHand-label 200-500 of your own reviews and measure accuracy first
Sarcasm and slang-heavy textA model trained on social media text, plus human review of low scores

Frequently asked questions

Is VADER still worth using?

Yes, as a baseline for short English text when you need speed and results you can explain word by word. Where context or other languages matter, a transformer model did clearly better in our test.

What do the VADER compound thresholds mean?

The compound score runs from -1 to +1. The VADER authors suggest positive at 0.05 or above, negative at -0.05 or below, and neutral in between. You can move the thresholds, but set them against labelled examples, not by feel.

Do I need a GPU for the Hugging Face pipeline?

No. The CPU build of torch scored our 24 reviews in seconds. For hundreds of thousands of reviews a GPU saves hours; pass device=0 to use it.

Which model works for Turkish reviews?

The nlptown model used here was not trained on Turkish. Search the Hugging Face hub for models whose card lists Turkish and product or review data, then test the best candidates on a labelled Turkish sample.

How do I measure accuracy on my own data?

Label a random sample of a few hundred reviews yourself and compare. Star ratings work as rough labels, as in this guide, but a three-star review often reads negative.

Can I use sentiment scores on scraped reviews?

Scoring is not the legal question; collecting is. Prefer an official API, check the site's terms and robots.txt, and keep personal data out.

Summary

Sentiment analysis in Python comes down to two choices. VADER from NLTK is small, fast and easy to explain, and it misreads sarcasm, slang, domain words and any language other than English. A Hugging Face model reads context and handles several languages at the cost of a large download and more compute; on our 24 reviews it matched the star rating on 19, against 14 for VADER. Whichever you choose, score per sentence or per review consistently, keep the confidence score, send disagreements to a person and measure against your own labelled sample. If your reviews come from public pages in several countries, see Proxynet proxies for collecting them at a polite rate.

Ask ChatGPTAsk Claude