A retailer reports its quarterly sales in six weeks. Before that, an analyst can already count how many products went out of stock on its website, how many store managers it is hiring and whether reviews of its new line are getting worse. None of these numbers come from the company's own filings, yet together they say something about the quarter. That kind of information is what the finance world calls alternative data.
This post explains what alternative data is and where it comes from: web data, card-transaction panels, satellite images, app usage and sentiment. It then follows a dataset from collection to decision, covers compliance and quality, shows where web scraping fits and compares building with buying. Nothing here is investment advice; the post is about data, not about which trades to make.
What is alternative data?
Alternative data is any dataset that helps explain a company, a sector or the economy but sits outside the traditional set of inputs. The traditional set is audited financial statements, regulatory filings, earnings calls, press releases, official statistics and market prices; everything else that carries a useful signal is "alternative".
The US Securities and Exchange Commission gives a working definition in a 2022 examination risk alert on MNPI compliance. Its examples include satellite and drone images of crop fields and parking lots, aggregate credit card transactions, social media and search data, phone geolocation data and email data from consumer apps. The same footnote adds that alternative data "does not necessarily contain MNPI", a point that comes back later in this post.
The term comes from investing, where hedge funds were the first large buyers; lenders, retailers, consultants and economists now use the same datasets.
What are the main types of alternative data?
The categories overlap, but most datasets fit one of these groups:
- Web data. Prices and stock levels on online shops, product catalogs, job postings, company headcount pages, store locators, app store rankings, real estate listings and reviews. It is public, fresh and cheaper to collect than other types.
- Transaction data. Aggregated and de-identified purchases from card issuers, payment processors or receipt apps. It shows spending by merchant or brand, usually through a panel of consumers rather than the whole market.
- Geospatial data. Satellite and aerial images (parking lots, oil tanks, crop fields, shipping) and foot-traffic counts derived from phone location signals.
- App and web usage. Download and active-user estimates for mobile apps, website traffic estimates and search volume trends.
- Sentiment and text data. Social media posts, news, product reviews, forum threads and earnings call transcripts turned into scores or topics. The CFA Institute notes in its article on using NLP on alternative data that much of this material is unstructured text and needs processing before it can be used.
How does alternative data become a usable signal?
The path from source to decision usually has six steps:
- Define the question. "Are sales at retailer X growing faster than last year?" is a question; "collect everything about retailer X" is not.
- Find and vet the source. Check who owns the data, how it was collected, under which terms, and whether you may use it for this purpose. For scraped data this means the site's terms and robots.txt; for a vendor it means a due diligence questionnaire.
- Collect. Scrape public pages, call an official API, receive files from a vendor or ingest a partner feed. Collection runs on a schedule, because a single snapshot shows nothing about change.
- Clean and join. Remove duplicates, fix types and currencies, map product names and store addresses to the right company and ticker.
- Build the indicator and test it. Turn rows into a series, for example a weekly count of discounted items, and compare it with past reported figures. A series that never tracked reported numbers is noise.
- Monitor. Sources change: a site redesigns its pages, a vendor's panel loses members, an app changes its store listing.
Steps 3 and 4 are an ETL job in all but name: extract, transform, load. What Is ETL? walks through such a pipeline with scraped data.
Alternative data types compared
| Type | What it measures | Typical collection | Freshness | Main risk |
|---|---|---|---|---|
| Web prices and stock | Pricing, discounts, availability | Scraping public product pages | Hours to days | Page changes break the scraper |
| Job postings | Hiring plans by role and location | Scraping career pages and job boards | Days | Duplicate and stale listings |
| Reviews and social posts | Sentiment, product problems | Scraping, official APIs | Hours | Bots, sarcasm, small samples |
| Card-transaction panels | Consumer spending by merchant | Licensed from vendors | Days to weeks | Panel bias, personal data |
| Satellite imagery | Physical activity (cars, tanks, crops) | Licensed imagery plus image analysis | Days | Clouds, cost, interpretation |
| Foot traffic | Visits to stores and venues | Phone location panels via vendors | Days | Consent and privacy |
| App usage estimates | Downloads, active users | Vendor models, app store data | Days | Opaque models, source quality |
Web data is the type a team can collect itself; card, location and app data almost always come from a vendor, which shapes both cost and compliance.
Is alternative data legal to use?
Using alternative data is legal in general; what matters is how a dataset was obtained and whether it carries information that should not be traded on.
Material nonpublic information (MNPI). In the US, investment advisers must keep written policies to prevent the misuse of material nonpublic information (Section 204A of the Investment Advisers Act). In the 2022 risk alert cited above, SEC examiners reported that some advisers used alternative data without policies that address the risk of receiving MNPI through it. The staff pointed to ad hoc diligence of data vendors, no assessment of the terms and legal obligations around collection (even after red flags appeared), diligence not applied to every source, and no process for repeating it when collection practices changed.
Misrepresenting where the data came from. In September 2021, the SEC charged App Annie and its co-founder with securities fraud. According to the SEC, the company told app makers their data would be aggregated and anonymized before use in its statistical model, but used non-aggregated data to adjust the estimates it sold to trading firms. The company agreed to a 10 million dollar penalty, and the SEC called it its first enforcement action against an alternative data provider.
Scraping and personal data. For web data, the questions are the ones any scraping project faces: are the pages public, do the site's terms allow automated access, does robots.txt ask you to stay out, and does the data include personal information covered by laws such as the GDPR or KVKK? Is Web Scraping Legal? covers the general picture, and Personal Data in Scraped Datasets covers how to reduce and protect personal fields.
Why is alternative data quality hard to judge?
Filings come with audits and standards; alternative data does not, so the buyer has to check it:
- Coverage. A card panel covers the people who use one issuer's cards or one receipt app, not the whole population. A scraper covers the pages it can reach, not the whole catalog.
- History. A signal needs years of history to be tested against reported results.
- Survivorship and changes. Stores close, sites change their product codes, vendors revise old data.
- Mapping. Brands, subsidiaries and store names have to be linked to the right company.
- Decay. When many firms buy the same dataset, the edge it gives shrinks. The CFA Institute article above makes the same point.
Where does web scraping fit?
Web scraping, the automated collection of public web pages, is how most teams build their own alternative data. It suits answers that are published online but spread over thousands of pages:
- Prices and promotions. Daily prices across retailers show discount depth and inflation in a category, the same technique covered in Competitor Price Tracking.
- Hiring. Job postings by role and city; LinkedIn Jobs by Country shows how listings differ by location.
- Reviews. Review counts and ratings over time, scored with methods like those in Sentiment Analysis in Python.
- Page changes. Updates to store locators, terms pages or pricing pages, tracked as in Website Change Monitoring.
Two points matter here. First, many sites show different prices, stock and listings by country, so a scraper must request pages from the right location, or the series mixes regions. Residential Proxy addresses let you choose the target country and city for each request. Second, a scraper that runs daily for years must be polite: respect robots.txt, keep request rates low, use an official API when one exists, and spread requests over time with Rotating Proxy settings rather than hitting one site in bursts. Proxies do not change what data you may collect; they only decide which network address the request comes from.
Should you build or buy alternative data?
| Factor | Build (collect yourself) | Buy (license from a vendor) |
|---|---|---|
| Control over method | Full; you know every step | Limited to what the vendor discloses |
| Time to first data | Weeks for a working scraper | Days after the contract |
| History | Starts from the day you begin | Often years of backfill |
| Ongoing work | Maintenance when sites change | Vendor due diligence, renewals |
| Compliance work | Your own review of sources and terms | Review of the vendor's sources and terms |
| Suitable data types | Public web data | Card, location, app, satellite data |
Most teams end up with both: vendor data for types they cannot collect, and their own scrapers for web data they need in a specific shape or country.
Where is alternative data used?
- Investment research. Analysts use web, card and app data to estimate sales before results are published; the finance proxy page covers data collection for financial firms.
- Market research. Consultants and brands size categories and follow competitors' ranges; see market research proxies.
- Retail pricing. Retailers compare their prices with competitors' every day, as described on the price monitoring page.
- Economic research. Online prices and job postings give earlier readings of inflation and labor demand than monthly statistics.
- Data products. Companies that build datasets for others rely on large-scale collection, the use case behind data scraping proxies.
Common mistakes with alternative data
- Starting from the dataset, not the question. A large dataset with no hypothesis produces charts, not answers.
- Skipping source diligence. Not knowing how a vendor collected its data is the exact gap the SEC alert describes.
- Mixing locations. Scraping a global site from one country and assuming the prices apply everywhere.
- Trusting a short backtest. A signal tested on one year may have fit by chance.
- Keeping personal data you do not need. Names, emails and user IDs rarely add signal and always add risk; drop them at collection time.
- Letting scrapers fail silently. A layout change that returns empty pages looks like "sales fell to zero" unless you monitor row counts.
- Treating the signal as permanent. Sources change and edges fade; review each dataset on a schedule.
Decision guide
| Need | Recommendation |
|---|---|
| Daily prices or stock from public shops | Build a scraper; request from the target country |
| Consumer spending by brand | License a transaction panel and review its sourcing |
| Physical activity at sites | License satellite or foot-traffic data |
| Hiring trends by company | Scrape career pages and job boards within their terms |
| Brand sentiment over time | Collect reviews or posts via APIs, score with NLP |
| Several years of history right now | Buy; a scraper cannot backfill the past |
| Signal nobody else has | Build, and document the method for compliance |
Frequently asked questions
What is the difference between alternative data and traditional data?
Traditional data is what companies and governments publish formally: statements, filings, earnings calls and official statistics. Alternative data is everything else that explains the same business, such as web prices or card panels. The line is about the source, not the format.
Is web scraping a source of alternative data?
Yes. Scraped web data (prices, product listings, job postings, reviews) is one of the most common types, and the one teams most often collect themselves. It should cover public pages only, follow site terms and robots.txt, and keep request rates low.
Who uses alternative data?
Hedge funds and asset managers came first. Private equity firms, lenders, retailers, consultancies and economists now use it too.
Can alternative data contain insider information?
It can. The SEC notes that alternative data does not necessarily contain material nonpublic information, but a dataset collected in breach of a confidentiality agreement or a site's terms could carry information that should not be traded on. That is why regulators expect firms to check how each source was collected before using it.
Is alternative data the same as big data?
No. Big data describes volume and processing methods; alternative data describes where information comes from. Some alternative datasets are huge, such as years of daily prices, while others are small, such as a weekly count of store openings.
How do I start with alternative data?
Pick one narrow question, find one source that could answer it, and collect enough history to test the answer against known results. For web data, a small scraper with clean storage and monitoring is enough to start; add vendors only when you need data you cannot collect.
Summary
Alternative data is information from outside the traditional set of filings, statements and official statistics: web prices, job posts, reviews, card panels, satellite images and app usage. It is useful when it answers a specific question earlier than traditional sources, after cleaning and testing against history. What separates a usable dataset from a risky one is source diligence: who collected the data, how, and under which terms. For the web data you collect yourself, public pages, polite rates and country-accurate requests matter most; Proxynet's proxy network provides the residential and rotating addresses for that collection.




