Scraping, field surveys, and panelists: What are the differences?
Fabrice Decroo
Consulting Director
August 16, 2026
In-store surveys, web scraping, and retailer panels each answer a different question: what the customer sees on the shelf, what is displayed online at a given moment, and what has actually been sold. Confusing them is like answering the wrong question with the right data.
Among large multichannel retailers, the proportion of prices that change each month rose from 15% to nearly 30% between 2008 and 2017—a pace that a weekly survey or monthly panel is structurally unable to keep up with (Alberto Cavallo, NBER).
“We’re already monitoring our competitors” means nothing until we’ve answered a specific question: What data, collected how, and how up-to-date is it? Three very different practices lie behind this statement—field surveys, web scraping, and retailer panels—and each answers a different question, with costs, timeliness, and reliability that are worlds apart. Confusing them isn’t just a methodological nuance—it means making a pricing decision based on the wrong data source. This guide compares the three methods point by point and provides the criteria for building an architecture that combines them rather than choosing just one by default.

Why These Three Methods Are Consistently Confused
In most retail organizations, “price monitoring” encompasses three distinct processes. Field data collection: A person observes and records a price, either in person or by manually checking it. Web scraping: a bot automatically queries web pages and organizes the displayed prices into actionable data— our definition article explains this mechanism in detail. Retailer panels: a third party aggregates data from point-of-sale systems or panelists, with an acknowledged time lag, to measure actual demand.
Treating these three sources as interchangeable comes at a high cost, in very concrete terms. A team that bases its weekly adjustments on panel data received several weeks late is always reacting after the market has already moved. Conversely, a team that bases a strategic repositioning on three days of scraping data mistakes statistical noise for a long-term trend, due to the lack of historical depth provided by a panel.
Field Survey
A person walks around or checks out the store, noting what they see: listed prices, as well as merchandising and actual out-of-stock items.
Web scraping
A bot automatically scans product pages and organizes the displayed prices into usable data.
Retailer panels
A third party aggregates point-of-sale or panelist data—with a time lag—to measure what was actually sold.
Choose based on the question asked
Each method answers a different question. Confusing them means answering the wrong question with the right data.
The following three sections explain, one by one, what each method actually captures—and, more importantly, what it does not capture.
Field Surveys: What They Capture, What They Don't Capture
Field data collection involves observing and recording—either in person or manually—the displayed price of a product at a retail location—such as a category manager walking through a competitor’s aisle, a commissioned investigator visiting a retail outlet, or an employee manually noting prices displayed online.
This is the only one of the three methods that captures what is actually happening on the shelf: product availability as seen by the customer—not the availability listed online—merchandising execution, consistency between the price advertised at the end cap and the price actually scanned at the register, and the visual impression a shopper has at the moment of decision. No web scraping tool, no matter how sophisticated, can detect an empty shelf or a poorly marked promotion, because a website may list a product as “available” even though it has been physically out of stock for two days.
This is the average out-of-stock rate observed on store shelves in a global sample of more than 71,000 consumers across 29 countries—a figure that no online data can reveal (Corsten & Gruen, Harvard Business Review, 2004).
Its limitation is structural, not technical: it doesn’t scale. A field survey of 200 SKUs across 5 retailers requires a fixed amount of man-hours; multiplying that by 10 categories and 20 retailers amounts to funding a dedicated full-time team, with costs that rise linearly while those of web scraping barely increase. For this reason, in most mature systems, fieldwork remains a targeted supplement—a one-time audit or sample-based quality control—rather than the foundation of day-to-day pricing management.
Web scraping: Freshness and granularity, but specific blind spots
Web scraping—whose three-layer technical process (collection, matching, and output) is detailed in our definitional article —is an automated process that queries public web pages at scheduled intervals and organizes the displayed prices into usable data. Its strength lies in two areas: timeliness—with the possibility of daily or even multiple daily data collection—and granularity on a product-by-product basis, without any scaling limitations imposed by human time constraints.
This freshness is not a matter of comfort; it is a necessity dictated by the market itself.
This is the proportion of prices that change each month at major multichannel retailers, which rose from 15% in 2008–2010 to nearly 30% in 2014–2017—a pace that a weekly survey or monthly panel is structurally unable to keep up with (Alberto Cavallo, NBER, Working Paper 25138, 2018).
Web scraping has specific blind spots, mirroring those of the physical store. It sees only what a website publishes—never the physical store itself, nor the price actually paid at the register after a discount not listed online, nor the actual availability on the shelves. It also depends entirely on the structure of the targeted site: a change in the template, enhanced anti-bot protection, or a page that switches to client-side dynamic rendering can interrupt data collection without anyone noticing, if the coverage rate isn’t continuously monitored.
Distributor Panels: The Truth Delayed
Retail panels—NielsenIQ, Kantar Worldpanel, and their industry-specific equivalents—do not measure the listed price; they measure the price actually paid, aggregated from checkout data or receipts reported by a sample of households. This is a difference in nature, not in degree, from web scraping. A panel tells us what was sold, at what price, and in what volume. Web scraping tells us what is displayed at a given moment, without any certainty as to what was actually purchased at that price.
The price of this reliability is a structural time lag. Data from a panel is collected, aggregated, deduplicated, and validated before being reported—a process that, depending on the provider and the market, can take several weeks.
stores worldwide, tracking 50 million products each month: the scale of the retail panels provides a comprehensive and reliable picture of actual demand, but at the cost of a monthly reporting cycle that is incompatible with responsive pricing management (NielsenIQ, Retail Measurement Services).
This time lag makes a panel unusable for reactive price adjustments—but that is precisely what makes it irreplaceable for validating a long-term trend. Data scraping may indicate that a competitor has been displaying a price 5% lower for the past three days; only a price panel can confirm, several weeks later, whether this displayed price difference resulted in an actual shift in volume or remained merely a display artifact with no measurable impact on demand.
A Detailed Comparison of the Three Methods
| Method | Freshness | Product Granularity | Cost | What Is Actually Measured |
|---|---|---|---|---|
| Field Survey | Punctual | Limited sample | High | Perception, merchandising, actual shelf availability |
| Web scraping | Real-time | Very detailed, item by item | Moderate | Price displayed online at this moment |
| Retailer panels | Several weeks | Grouped by category / reference | High | Price actually paid, volume sold |
To put it this way: the “what is actually measured” column is the most important of the four. A scraped listed price, a price paid as measured by a panel, and a perception observed in the field are not three more or less accurate versions of the same information—they are three different pieces of information that answer three different questions. Confusing the columns is like answering the wrong question with the right data.
How to Combine Them in a Hybrid Architecture Based on Your Maturity Level
The question is never “scraping, fieldwork, or panel,” but rather in what order and for what purpose to combine them. The maturity of a price monitoring system is built in stages, not through a binary choice made once and for all.
One-time assignment + spreadsheet
Initial manual data entries compiled by hand, often in a spreadsheet. Works on a small scale; cannot handle a larger workload.
Single-source scraping for day-to-day management
Automated scraping takes over for the most volatile categories, though product matching is still rudimentary.
Multi-source scraping + targeted fieldwork for quality control
Matching is becoming more reliable, coverage is expanding, and an occasional field survey verifies what data scraping cannot detect.
Governed architecture, organized by category
Data scraping, field research, and panel surveys coexist, each assigned to the decision it best informs, with a governance structure that resolves discrepancies.
At each stage, the question that should guide the decision isn’t “which method is best” in absolute terms, but rather which benchmarks and categories justify the investment—a detailed prioritization outlined in our article on the products you should really be monitoring. Many organizations get stuck at the first stage, using a system cobbled together in a spreadsheet, for reasons documented in our article on the limitations of Excel for price monitoring.
At Booper
This is exactly the path taken by the Barbotteau Group, the leading retailer in the French Caribbean with more than 70 companies and brands: a management approach that historically relied on Excel, then moved to a rules-based system, followed by hybrid AI, and finally full automation—each step tailored to the organization’s specific needs. After this data is collected—regardless of the source—the GENIUS Link module automatically matches SKUs using NLP, with a confidence score for each product, and GENIUS Monitoring centralizes alerts so that each source informs a decision rather than just adding to another spreadsheet.
The right question isn't "which one should I choose?"
None of the three methods is universally superior to the other two—each answers a question that the other two cannot ask in its place. Field research answers the question, “What does the customer see and experience on the shelf?” Web scraping answers, “At what price is my competitor currently listing this product?” The panel survey answers, “What actually sold, and at what price?” Find out how Booper structures this comprehensive architecture on our price tracking & web scraping page.
A Checklist to Review Before Evaluating Your Price Monitoring Architecture
- Do you know which decision each data source is supposed to inform, or do you use them interchangeably?
- Does your scraping process capture the up-to-date information needed for your most volatile categories?
- Does a field survey—even a one-time one—verify what web scraping can't detect (actual out-of-stock items, merchandising)?
- Is your panel data used to validate a trend or for day-to-day management—a common point of confusion?
- Does your system combine these sources, or does it use only one by default due to a lack of well-thought-out architecture?
FAQ
The questions we are most frequently asked before getting started.
Field data collection involves observing and recording—either in person or through manual data entry—the displayed price of a product at a retail location—such as a category manager walking through a competitor’s aisle or a surveyor visiting a retail outlet. It is the only method that captures what is actually happening on the shelf: product availability as seen by the customer, the execution of merchandising, and consistency between the price advertised at the endcap and the price scanned at the register.
Web scraping, on the other hand, is an automated process that queries public web pages at scheduled intervals and organizes the displayed prices into usable data— without the scalability limitations imposed by human time, and with data freshness that can be as up-to-date as real time.
The flip side is a blind spot that mirrors that of the store floor: scraping never detects an empty shelf, a poorly marked promotion, or online availability that doesn’t reflect the physical reality of the store. A website may list a product as “available” even though it has been out of stock for two days.
For this reason, in mature systems, fieldwork remains a targeted supplement—such as ad hoc audits or sample-based quality control—rather than the foundation of day-to-day management, due to the inability to scale up: a survey of 200 SKUs across five retail chains already requires a minimum amount of staff time that cannot be reduced.
No. A retail panel does not measure the listed price, but rather the price actually paid, aggregated from checkout data or receipts reported by a sample of households—a difference in nature, not in degree, from web scraping.
The price of this reliability is a structural time lag: data is collected, aggregated, deduplicated, and validated before being reported—a process that, depending on the provider and market, can take several weeks. NielsenIQ, for example, tracks approximately 900,000 stores worldwide for 50 million products each month—a massive scale, but at a monthly cadence that is incompatible with responsive pricing management.
This time lag makes a panel unusable for day-to-day price adjustments, but that is precisely what makes it irreplaceable for confirming a long-term trend: price scraping may indicate that a competitor has been offering a price 5% lower for the past three days, but only a panel can confirm—several weeks later—whether that difference has resulted in an actual shift in sales volume.
A comprehensive price monitoring system therefore requires both approaches: scraping for day-to-day responsiveness, and a panel for in-depth validation.
No, it captures the price displayed at a given moment—not necessarily the price actually paid. A discount applied only at the register, a negotiated price, or a loyalty program benefit not reflected online are not visible through web scraping.
That is precisely the role of retailer panels: they measure the actual price paid and the volume sold, aggregated from point-of-sale data or consumer panels, whereas web scraping only captures what a website publishes.
Web scraping also has blind spots that depend on the technical structure of the targeted site: a change in the template, enhanced anti-bot protection, or a page that switches to client-side dynamic rendering can interrupt data collection without anyone noticing if the coverage rate isn't continuously monitored.
That is why the listed scrap price, the price paid as measured by a panel, and field observations are not three versions of the same information with varying degrees of accuracy: they are three different pieces of information that answer three different questions.
Combine them. Data scraping supports day-to-day pricing management thanks to its up-to-date nature; field staff periodically check for what the scraping process cannot detect (actual out-of-stock situations, merchandising); and the panel verifies that the observed discrepancies actually translate into real volume.
The right balance is achieved through stages of maturity rather than a one-time, binary choice: starting with ad hoc fieldwork compiled in a spreadsheet, then single-source scraping of the most volatile categories, followed by multi-source scraping with fieldwork for quality control, and finally a governed architecture where all three sources coexist, each assigned to the decision it best informs.
The question that should guide the decision is never “which method is best” in absolute terms, but rather which benchmarks and categories justify investing in each one.
This is exactly the path taken by the Barbotteau Group, a retail leader in the French Caribbean with more than 70 companies and brands: a management approach that historically relied on Excel, then moved to a rules-based system, followed by hybrid AI, and finally full automation—each step tailored to the organization’s specific needs.
Because their data comes from the collection, aggregation, deduplication, and validation of point-of-sale data or large-scale consumer panel data—a process that, to remain statistically reliable, cannot be instantaneous.
Unlike web scraping, which structures data collected directly from a web page, a panel must gather data from thousands—or even millions—of retail locations or households before it can produce a representative figure. NielsenIQ, for example, tracks approximately 900,000 stores worldwide for 50 million products each month—a scale that explains why the data is reported monthly rather than daily.
This delay is not a flaw to be corrected; it is the trade-off for the reliability we seek: a panel provides data on what was sold, at what price, and in what volume, with a level of statistical rigor that an instantaneous data collection could not guarantee.
This is what makes a panel suitable for validating underlying trends, but structurally unsuitable for day-to-day, reactive pricing management—a role that is better fulfilled by web scraping.
Yes, within a targeted scope. Web scraping never captures an empty shelf, a poorly labeled promotion, or the actual customer experience on the sales floor —elements that only an on-site survey, even an occasional one, can objectively document.
According to a Harvard Business Review study, the average in-store out-of-stock rate observed across a large international sample is 8.3% —a figure that no online data can reveal, since a website may show a product as available even though it has actually been out of stock in stores for several days.
Fieldwork therefore retains its place, no longer as the foundation of day-to-day management—its limitation in terms of scalability remains structural—but as a targeted supplement: an ad hoc audit, a sample-based quality control check, designed to verify what scraping is structurally unable to capture.
In a mature price monitoring architecture, these three sources (field data, web scraping, and panel data) coexist, with each one assigned to the decision it best informs, rather than being treated as interchangeable.
Also in this series
- What is price tracking or web scraping in retail?
- Competitors' Prices: Which Products Should You Really Keep an Eye On?
- The limitations of Excel for tracking competitor prices
Sources: NBER, Alberto Cavallo, “More Amazon Effects: Online Competition and Pricing Behaviors,” Working Paper 25138, 2018 · Corsten & Gruen, “Stock-Outs Cause Walkouts,” Harvard Business Review, May 2004 · NielsenIQ, Retail Measurement Services.
Seven criteria distinguish a competitor pricing monitoring tool that simply generates a table of price differences from one that actually drives decisions: coverage, recency, product matching, alerts, governance, integration, and compliance. The listed cost is only part of the total cost: manual reclassification, maintenance of in-house development, and the opportunity cost of a poorly informed decision often outweigh the subscription fee.
A comprehensive competitor pricing monitoring system is built on five inseparable components: data collection, matching, alerts, reporting, and governance—if even one of these components is missing, the system becomes ineffective. The retail sector revises its prices more frequently than any other (ranging from monthly to daily, depending on the category), which requires a system capable of keeping pace with this frequency.
Many organizations receive a report on competitor price gaps every morning—but few have a genuine strategy. The difference lies in three questions asked before implementing the system: Why collect this data? What exactly should be tracked? And what decisions should be made once a price gap is identified?
Key point: Key value items (KVI) —the products for which customers remember the price—typically account for 15 to 25 percent of a category’s sales. Focusing monitoring efforts on this small core group is more cost-effective than trying to track everything with the same intensity.
