Resources / Glossary
The web data glossary.
288 terms across web scraping, anti-bot, data quality, delivery, pricing, the digital shelf, AI data and compliance — defined plainly by the people who have run web data in production since 2012.
A
Acceptance test
A defined check a dataset must pass before it is accepted, such as row counts, fill rates and value ranges.
Agentic browsing
An AI agent operating a web browser to navigate, click and fill in forms to complete a task.
AI agent
Software that uses a language model to plan and take actions, such as calling tools, browsing or running code, towards a goal.
AI crawler
A crawler operated to collect web content for AI training, search or agents, usually identified by its user agent.
AI extraction
Using machine learning models, including large language models, to identify and extract fields from pages or documents.
AI QA
Automated checks that compare extracted records with the source page and the schema before delivery.
Alerting
Notifying people or systems when a defined condition occurs, such as a price change, stock-out or failed delivery.
Alternative data
Non-traditional data used by investors, such as web prices, job postings, app usage or satellite images.
Anomaly detection
Identifying values or patterns that deviate from what is expected, such as a sudden price drop or a collapse in row count.
Anti-bot system
Technology that websites use to detect and block automated traffic, using signals such as IP reputation, fingerprints, behaviour and challenges.
Aperture
Import.io's pricing intelligence and MAP monitoring application, with AI product matching and dated evidence.
API
A defined way for software to request data or actions from another system.
API key
A secret token that identifies and authorises a client calling an API.
ASIN
Amazon's 10-character identifier for a product listing.
Assortment
The set of products a retailer offers in a category.
Assortment gap
A product, or type of product, that competitors carry and you don't, or the reverse.
Audit trail
A chronological record of actions and changes that can be reviewed later.
Authenticated collection
Collecting data from pages that require signing in, using credentials the account holder has authorised for that use.
Automated extraction
Extraction that identifies records and fields on a page without a person writing selectors, using page structure, heuristics or machine learning.
Availability
Whether a product can be bought now, and sometimes where and how quickly.
B
Backfill
Collecting or reprocessing historical data to fill a gap or populate a new field over past periods.
Baseline price
The normal, non-promotional price of a product, used as the reference for measuring discounts.
Benchmark contamination
When evaluation examples leak into a model's training data and inflate its scores.
Benchmarking
Comparing performance, prices or practices against competitors or a standard.
Boilerplate removal
Stripping navigation, footers, adverts and other repeated elements so only a page's main content remains.
Bot detection
Classifying traffic as human or automated from network, browser and behavioural signals.
Brand protection
Protecting a brand's pricing, reputation and intellectual property online, including monitoring sellers, counterfeits and misuse.
Breadcrumbs
The navigation trail on a page that shows where it sits in a site's hierarchy, such as Home › Tools › Drills.
Browser fingerprinting
Identifying a client from the combination of its browser, device, fonts, screen and other attributes.
Buy Box
The default offer on a marketplace product page, which most shoppers buy from.
Buy Box win rate
The share of time or observations in which a seller holds the Buy Box on a listing.
C
Cadence
How often data is collected and delivered.
Canonical URL
The preferred address for a page, often declared with a rel=canonical tag, when the same content is reachable at several URLs.
CAPTCHA
A challenge meant to be easy for people and hard for software, such as selecting images or ticking a box.
Catalogue coverage
The share of a brand's or competitor's catalogue that a retailer lists.
Category management
Managing a product category as a business unit: range, pricing, promotion and space.
CCPA and CPRA
California laws that give residents rights over personal information collected about them.
Chaining
Linking extractors so the output of one, such as product URLs from a listings page, becomes the input of the next, such as a details-page extractor.
Change detection
Identifying what changed on a page or in a dataset since the last observation.
Chunking
Splitting documents into smaller passages for embedding and retrieval.
Clearance
Pricing intended to sell remaining stock, usually of discontinued or seasonal items.
Competitive intelligence
Gathering and analysing information about competitors' products, prices, moves and positioning.
Completeness
How much of the expected data is present: records against the known universe, and fields within each record.
Compliance filters
Rules applied during collection or processing to exclude disallowed sources, fields or personal data.
Concurrency
The number of requests a system makes at the same time.
Consistency
Whether the same thing is represented the same way across records, sources and time.
Content compliance
Whether a product page's content matches the brand's approved titles, images, descriptions and specifications.
Content hash
A fixed-length fingerprint computed from content, used to detect changes and duplicates.
Cookies
Small pieces of data a site stores in the client to keep sessions, preferences and tracking state.
Copyright
Legal protection for original creative works, such as text, images and code.
Counterfeit
A fake product sold as genuine.
Coverage
The share of target sources, pages or entities that a dataset actually contains.
Crawl
A run of a web crawler across a site or set of sites, discovering and fetching pages according to rules on scope, depth and rate.
Crawl budget
The number of requests a crawl may make in a period, set by cost, politeness or a site's tolerance.
Crawl depth
How many link hops away from the seed a crawler will follow.
Crawl politeness
Practices that limit a crawler's load on sites, such as rate limits, off-peak scheduling and honouring robots.txt.
Crawl run
In Import.io, one execution of an extractor or crawl over its inputs, with its own results, counts and status.
CSS selector
A pattern that identifies elements in an HTML document by tag, class, id or attribute, commonly used to target the fields a scraper extracts.
CSV
A plain-text table format with one record per line and fields separated by commas.
Currency normalisation
Converting prices into a common currency and format, recording the exchange rate and date used.
D
Dashboard
A visual display of key metrics, updated from underlying data.
Data aggregation
Combining data from many sources into a unified dataset or summary.
Database right
In the EU and UK, protection for databases that required substantial investment, separate from copyright.
Datacenter proxy
A proxy that uses IP addresses belonging to cloud or hosting providers.
Data drift
A change in the statistical properties of data over time, such as a shift in price or category distributions.
Data enrichment
Adding attributes to existing records from other sources, such as appending web data to a product catalogue.
Data extraction
The step that turns raw web content into structured records with defined fields and types, ready to load, query or analyse.
Data feed
A recurring delivery of data from a source to a destination on a schedule.
Data governance
The policies, roles and controls that decide how data is collected, stored, used and retired.
Data labelling
Adding tags or answers to data so it can train or evaluate models.
Data lake
Storage that holds raw and processed data in open file formats, often on object storage.
Data lineage
The record of where data came from and every step that transformed it on the way.
Data minimisation
Collecting and keeping only the personal data needed for a defined purpose.
Data observability
Visibility into the health of data pipelines through metrics, logs, lineage and quality checks.
Data pipeline
The sequence of steps that moves data from sources, through processing, to destinations.
Data processing agreement (DPA)
A contract governing how a processor handles personal data on behalf of a controller.
Data quality
How well data fits its purpose, usually measured as accuracy, completeness, freshness, consistency and uniqueness.
Data report
A structured summary of data prepared for a specific audience or decision.
Data residency
Requirements on where data is stored and processed geographically.
Data retention
Rules on how long data is kept and when it is deleted.
Data warehouse
A database built for analytics across large volumes of structured data, such as Snowflake, BigQuery or Redshift.
Deduplication
Removing repeated records, whether exact copies or near-duplicates that refer to the same entity or document.
Delivery manifest
A file that accompanies a data delivery and lists its contents, counts, schema version, checks and provenance.
Details page
A page dedicated to a single item, such as a product, property or job, carrying its full set of attributes.
Digital shelf
Everywhere a product appears online where shoppers find, compare and buy it: retailer sites, marketplaces and search.
Digital shelf analytics
Measuring and improving a product's online presence: availability, price, content, search rank, reviews and share of shelf.
Document parsing
Extracting text, tables and fields from documents such as PDFs, filings and spreadsheets.
DOM (Document Object Model)
The tree of elements a browser builds from HTML, which scripts modify and selectors query.
Dynamic pricing
Changing prices frequently in response to demand, competition, stock levels or time.
E
EAN
A 13-digit GTIN used internationally.
Embeddings
Numeric vectors that represent the meaning of text, images or other data, so similar items sit close together.
Entity resolution
Determining which records across sources refer to the same real-world entity, such as a product, company or property.
ETL (extract, transform, load)
Extracting data from sources, transforming it and loading it into a destination. ELT loads first and transforms inside the warehouse.
Evaluation set
A held-out set of examples with known answers, used to measure a model's performance.
Exact match
Two listings of the identical product: same brand, model, size and pack.
External data
Data from outside an organisation, such as market, competitor, public-record or web data.
Extractor
In Import.io, a configured set of rules that turns a website's pages into structured records, built by point-and-click, AI assistance or engineers.
F
Feed delivery
Getting data from the extraction system to where it is used: files, warehouse tables, APIs or webhooks.
Field accuracy
The share of extracted values that match the source when checked, typically against a human-labelled sample.
Fill rate
The share of records in which a field has a value.
Fine-tuning
Further training of a pretrained model on a smaller, targeted dataset to specialise it.
First-party (1P) retail
An arrangement in which a marketplace buys stock from a brand and sells it directly.
First-party data
Data a company collects directly from its own customers and operations.
Freshness
How recent data is, measured as the time between a source changing and the data reflecting it.
Fulfilment
How an order is stored, packed and shipped, for example by the seller or by the marketplace.
Full crawl
A crawl that refetches every page in scope, whether or not it changed.
Fully managed service
An arrangement in which the provider builds, runs, quality-checks and maintains data collection and delivers the data under an agreement.
G
GDPR
The EU regulation governing the processing of personal data about people in the EU.
Geo-targeting
Sending requests from a specific country, region or city so a site returns its local content, prices and availability.
Governance
The framework of decisions, accountability and controls around a data program.
GraphQL
A query language for APIs that lets clients request exactly the fields they need.
Grey market
Genuine products sold through channels the brand didn't authorise, often imported from other markets.
Grounding
Tying a model's output to specific source material so its claims can be checked.
Ground truth
Data verified by people against the source, used to measure extraction accuracy.
GTIN
A standard product identifier issued under the GS1 system, including UPC (12 digits) and EAN (13 digits).
H
Hallucination
A model producing fluent output that is false or unsupported by its inputs.
Headless browser
A real web browser run without a visible window and controlled by code to load pages, execute JavaScript and interact like a user.
Historical data
Past observations kept over time, such as daily prices over several years.
Honeypot
A hidden link or form field placed to catch automated clients that a person would never interact with.
HTML
The markup language that structures web pages into elements such as headings, tables, links and images.
HTTP headers
Metadata sent with web requests and responses, such as user agent, cookies, language and caching directives.
HTTP status codes
Three-digit codes a server returns to report the result of a request, such as 200 OK, 403 Forbidden, 404 Not Found or 429 Too Many Requests.
Human in the loop
A process in which people review or correct automated output at defined points.
I
Image URL
The web address of an image file, such as a product photo.
Incremental crawl
A crawl that fetches only pages that are new or have changed since the last run.
Infinite scroll
A pattern where more results load automatically as the user scrolls, instead of through page links.
Ingestion
Bringing data into a storage or processing system from outside sources.
Insights
Conclusions drawn from data that change a decision.
In-stock rate
The share of tracked products that are available across retailers, stores or time.
IP block
A site refusing requests from an IP address or range it has flagged.
IP rotation
Changing the IP address used for requests across a pool, per request or per session.
ISP proxy
A proxy on IP addresses registered to internet service providers but hosted in data centres.
J
JavaScript challenge
A script a site serves to check that the client is a real browser before showing content.
JavaScript rendering
Executing a page's scripts, usually in a headless browser, so content loaded after the initial HTML becomes available to extract.
JSON
A lightweight text format for structured data built from key-value pairs and arrays.
JSON-LD
A JSON-based format for embedding linked data, most often Schema.org markup, in a web page.
JSON Lines
A format with one JSON object per line, easy to stream, split and append.
K
Key value item (KVI)
A product whose price shoppers remember and use to judge a store's overall price level.
Keyword ranking
The position at which a product appears for a given search term on a retailer or marketplace.
KPI
A measure chosen to track progress against a goal.
L
Landed price
The full price to the buyer, including shipping, fees and taxes.
Language identification
Detecting which language a document or passage is written in.
Lawful basis
Under the GDPR, the legal ground for processing personal data, such as consent or legitimate interests.
Lazy loading
Deferring content such as images, prices or reviews until it is about to be viewed.
Licence signal
Information on a page that indicates its licence or usage terms, such as a Creative Commons tag.
Link extraction
Collecting URLs from a page to discover further pages to crawl or extract.
Listings page
A page that shows many items at once, such as a category, search result or feed, each linking to a details page.
List price
The price a manufacturer suggests or a retailer shows as the reference, before discounts.
llms.txt
A proposed file that sites publish to give language models a curated guide to their content.
Localisation
The adaptation of a site's content, language, currency and assortment to a market.
Loyalty price
A price available only to members of a retailer's loyalty scheme.
M
MAP (minimum advertised price)
A manufacturer policy setting the lowest price at which resellers may advertise a product, common in the US.
MAP policy
A manufacturer's written rules on minimum advertised prices, including scope, exceptions and enforcement steps.
MAP violation
An advertised price below the minimum set in a manufacturer's MAP policy.
Markdown
A permanent or long-term price reduction, often made to clear stock.
Market intelligence
Information about a market — its size, players, prices and trends — gathered to inform strategy.
Marketplace
A platform where many third-party sellers list products, such as Amazon, eBay or Walmart Marketplace.
Match confidence score
A score expressing how likely it is that two listings are the same product.
MCP (Model Context Protocol)
An open protocol that lets AI applications connect to external tools and data sources through standard servers.
Mobile proxy
A proxy that routes traffic through mobile carrier networks.
Monitoring
Continuous checking of sources, runs and outputs for failures, anomalies and changes.
MPN
The identifier a manufacturer assigns to a product model.
MSRP
The retail price a manufacturer recommends for a product, a term used mainly in the US.
Multibuy
A promotion that rewards buying more than one unit, such as buy one get one free or three for two.
N
National brand
A manufacturer's brand sold through many retailers.
Near-duplicate detection
Finding documents or records that are almost identical, using techniques such as shingling and locality-sensitive hashing.
Normalisation
Converting values to a consistent form: units, currencies, formats, names and categories.
O
Object storage
Cloud storage that organises files as objects in buckets, such as Amazon S3.
OCR
Converting text in images or scanned documents into machine-readable characters.
Offer price
The price at which a product is currently offered to buyers, after any visible discounts.
Omnichannel
Selling across online and physical channels as one connected experience.
Organic rank
A product's position in results, excluding paid placements.
Out of stock
A product that is listed but cannot be bought because no inventory is available.
P
Pagination
Splitting a list of results across multiple pages, navigated by page numbers, next links, offsets or cursors.
Parquet
A columnar file format that compresses well and stores data types, widely used in data lakes and warehouses.
Personal data (PII)
Information relating to an identifiable person, such as names, contact details or online identifiers.
PII redaction
Detecting and removing personal information such as names, email addresses and phone numbers from data.
Point-and-click extraction
Building an extractor by clicking on the data you want in a visual interface instead of writing code.
Point-in-time data
Data recorded as it was known at each moment, without later revisions, so backtests avoid look-ahead bias.
Pretraining
The first, broad phase of training a model on large general datasets.
Price dispersion
The spread of prices for the same product across sellers or channels at one moment.
Price elasticity
How much demand for a product changes when its price changes.
Price index
A single number comparing your prices with a competitor's or the market's across a basket of matched products, often with 100 as parity.
Price intelligence
Collecting and analysing competitors' prices, promotions and availability to inform pricing decisions.
Price monitoring
Regularly capturing prices for defined products across competitors and channels.
Price optimisation
Setting prices using models of demand, elasticity, cost and competition to meet goals such as margin or volume.
Price parity
Keeping the same price for a product across channels, regions or sellers.
Price perception
How expensive customers believe a retailer is, which can differ from its actual average prices.
Price position
Where a product's price sits relative to competitors: cheapest, at parity, or at a premium by some margin.
Private label
Products sold under a retailer's own brand.
Product content
The titles, descriptions, images, specifications and other material on a product page.
Product data
Information that describes products: identifiers, titles, attributes, images, prices, availability and reviews.
Product matching
Identifying the same product across retailers or sellers, using identifiers, attributes, text and images.
Product Q&A
Shopper questions and seller or community answers shown on product pages.
Product taxonomy
A hierarchy that classifies products into categories and subcategories.
Promotion
A temporary offer that changes the price or value of a product, such as a discount, multibuy, coupon or loyalty price.
Provenance
Evidence of a record's origin: source URL, capture time, collection method and version.
Proxy
An intermediary server that forwards requests, so the target sees the proxy's IP address instead of the client's.
Proxy pool
The set of IP addresses available to a collection system for routing requests.
Publicly available data
Information that can be accessed on the web without logging in or paying.
Q
Quality filtering
Removing low-quality documents from a corpus, such as spam, boilerplate or machine-generated text.
Quality score
A composite measure summarising a record's or dataset's quality across checks such as validity, completeness and consistency.
Query
In Import.io, one successful request that returns data. Self-service plans are billed on successful queries.
R
RAG (retrieval-augmented generation)
A pattern in which a model retrieves relevant documents at query time and uses them to ground its answer.
Ranking
The order in which products appear in search results, categories or bestseller lists.
Rate limit
A cap on how many requests a client may make in a period, enforced by the site or chosen by the collector.
Rate parity
In travel, offering the same room rate or fare across booking channels.
Recrawl frequency
How often a source is fetched again, such as hourly, daily or weekly.
Repricing
Automatically changing a price according to rules or models, often in reaction to competitor prices.
Resale price maintenance
Arrangements that fix or impose minimum resale prices on resellers, which many jurisdictions restrict.
Residential proxy
A proxy that routes traffic through IP addresses assigned to home internet connections.
REST API
An API that exposes resources through standard HTTP methods such as GET and POST, usually returning JSON.
Retail media
Advertising sold by retailers and marketplaces on their own sites and apps.
Retries and backoff
Repeating failed requests after increasing delays, to recover from transient errors without overloading the target.
Retrieval freshness
How current the documents in a retrieval index are relative to the live web.
Reviews and ratings
Customer feedback on products: star ratings, written reviews, review counts and helpfulness votes.
Review velocity
The rate at which new reviews arrive for a product.
robots.txt
A file at the root of a site that tells automated agents which paths they may or may not crawl.
RRP
The retail price a supplier recommends, a term common in the UK and EU.
RRP monitoring
Tracking how resellers' prices compare with a recommended retail price.
S
Schema
The definition of a dataset's fields, types, allowed values and rules.
Schema contract
A written agreement on a dataset's fields, types, rules and acceptance tests between the producer and the consumer.
Schema drift
Unplanned changes to a dataset's structure over time, such as renamed, added or retyped fields.
Schema.org markup
A shared vocabulary that sites embed in pages, often as JSON-LD, to describe products, prices, reviews and organisations for search engines.
SDK
A library that wraps an API for a particular programming language, handling authentication, retries and data types.
Seed URL
A starting address from which a crawl begins discovering pages.
Selector drift
The mismatch that develops between an extractor's selectors and a site's markup when the site changes.
Self-healing extractor
An extractor that detects when a site change breaks it and adapts, automatically or with review, to restore correct output.
Self-service plan
An Import.io subscription in which your team builds and runs extractors on the platform, with published pricing and query allowances.
Seller identification
Resolving marketplace seller names and IDs to the businesses behind them.
Semi-structured data
Data with some organisation but no fixed schema, such as HTML, XML or JSON whose fields vary between records.
Sentiment analysis
Classifying text such as reviews or articles as positive, negative or neutral, often by topic.
Session
A sequence of requests that a site treats as one visit, tied together by cookies or tokens.
SFTP
A protocol for transferring files over an encrypted connection.
Shrinkflation
Reducing a product's size or quantity while keeping its price the same or similar.
Similar match
A listing that is not identical but is a close substitute, such as a competing private-label item.
Single-page application (SPA)
A website that loads one page and then updates its content with JavaScript instead of loading new pages from the server.
SKU
A retailer's or seller's own identifier for a product or variant.
SLA
A contract term defining measurable commitments, such as coverage, freshness, accuracy and response time, with remedies if they are missed.
Snapshot
A stored copy of a page or dataset exactly as it was at a specific moment.
Sponsored listing
A paid placement in a retailer's or marketplace's search results.
Stock level
The quantity of a product available, where a site reveals it through counts or thresholds.
Store-level pricing
Prices that vary by store, postcode or region for the same product at the same retailer.
Strike-through price
A higher reference price shown crossed out next to the current price.
Structured data
Information organised into defined fields and types, such as rows in a table or keys in a JSON object.
Synthetic data
Artificially generated data that mimics the properties of real data.
T
Table extraction
Turning HTML or document tables into rows and columns with the correct headers and data types.
Terms of service
The conditions a website sets for its use.
Third-party data
Data collected by one organisation and sold or shared to others.
Third-party seller (3P)
An independent business selling on a marketplace, as opposed to the marketplace selling directly.
Time series
Observations of the same measure recorded at successive points in time.
TLS fingerprinting
Identifying clients by how they negotiate encrypted connections, such as cipher order and extensions.
Tokens
The units of text a language model processes, typically whole words or word pieces.
Tool calling
A model's ability to request that an external function or API be run, and to use the result.
Training (an extractor)
Teaching an extractor which data to capture by selecting examples on a page, so it can find the same fields on similar pages.
Training corpus
A large, curated collection of documents assembled to train or fine-tune a model.
Training data
The examples a machine learning model learns from.
Transformation
Changing data after extraction — cleaning, typing, splitting, joining and deriving fields — to fit the target schema.
U
Unauthorised seller
A seller offering a brand's products without being an approved reseller.
Unblocker
A service that handles proxies, fingerprints, challenges and retries to return a page from a protected site.
Unit price
The price per standard measure, such as per kilogram or per 100 ml.
Unstructured data
Content without a predefined data model, such as free text, images, audio or PDFs.
UPC
A 12-digit GTIN used mainly in North America.
Uptime
The share of time a service is operational.
URL
The address of a resource on the web, made up of a scheme, host, path and optional query parameters.
URL frontier
The queue of discovered URLs a crawler has yet to fetch, with rules for priority and deduplication.
URL normalisation
Rewriting URLs to a standard form — ordering parameters, removing tracking codes, fixing case — so duplicates are recognised.
User agent
The request header in which a client identifies its software, such as a browser name and version.
User behaviour data
Data describing how people interact with sites and apps, such as clicks, searches and sessions.
V
Variant
A version of a product that differs by an attribute such as size, colour or pack.
Vector database
A database built to store embeddings and find nearest neighbours quickly.
Visibility
How easily shoppers come across a product online, combining rank, placement, availability and content.
W
Web application firewall (WAF)
A security layer in front of a site that filters requests by rules and threat signals.
Web archive
A collection of web pages stored over time, such as the Internet Archive's Wayback Machine.
Web crawler
A program that discovers pages by following links from a set of starting URLs and fetching each page it finds.
Webhook
An HTTP callback that notifies another system when an event happens, such as a delivery completing.
Web Scraper MCP
Import.io's MCP server, which gives AI agents web data extraction as tools in Claude, Cursor and other MCP clients.
Web scraping
Automated collection of information from websites by software that requests pages and pulls out the content it needs.
Web scraping service
Buying web data as a service: the provider owns extraction and maintenance and delivers data on a schedule.
Web search API
An API that returns search results, and often page contents, programmatically, commonly used by AI applications.
X
XML
A markup format for structured documents and data exchange, used for feeds, sitemaps and many older APIs.
XML sitemap
A file listing a site's URLs, often with last-modified dates, published to help search engines discover pages.
XPath
A query language for selecting nodes in an HTML or XML document by path, position and conditions.
Y
Yield
The share of fetched pages or requests that produce valid records.
Z
Zero-party data
Data a customer intentionally and proactively shares with a brand, such as preferences or intentions.