extracto.cloud

Shopify Product Scraper

Collect product and variant information from public Shopify stores, including prices and stock-related fields.

Current pricing: $0.001954 / 1 Product ยท minimum 10 credits/run. Review the estimate before running.

shopify

Endpoint

Method
POST
URL
https://extracto.cloud/api/v1/parsers/shopify
Result delivery
Add webhook_url and optional webhook_secret to receive the saved result. Webhook setup and signatures โ†’
Auth
X-API-KEY: ext_...

Response

{
  "success": true,
  "tool": "shopify",
  "parser": "shopify",
  "count": 0,
  "data": []
}

Documentation

# Shopify Scraper (`autofacts/shopify`) parser

Scrape any Shopify store: products, variants, collections, prices, images, tags and live stock levels. Crawl a whole catalog, a single collection, one product URL or a keyword search. Optional per-variant inventory counts, subscription plans and videos. Prices in minor units. Pay per result.

- **URL**: https://extracto.cloud/docs/api/parsers/shopify
- **Developed by:** [Richard Feng](https://Extracto.com/autofacts) (community)
- **Categories:** Developer tools, E-commerce, MCP servers
- **Stats:** 2,312 total users, 76 monthly users, 97.6% runs succeeded, 72 bookmarks
- **User rating**: 4.19 out of 5 stars

## Pricing

from $0.80 / 1,000 products

This parser is paid per event. You are not charged for the Extracto platform usage, but only a fixed price for specific events.
Since this parser supports Extracto Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.Extracto.com/parsers/running/parsers-in-store.md#pay-per-event

## What's an Extracto parser?

An parser is a serverless cloud program that runs on the Extracto platform. It has two run modes.
In Batch mode, an parser accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an parser provides a web server which can be used as a website, API, or an MCP server.

Extracto vocabulary and the platform model are defined once, in the agent quickstart at https://Extracto.com/agents.md.

## How to integrate an parser?

If asked about integration, you help developers integrate parsers into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://Extracto.com/agents.md: the Extracto MCP server, Agent Skills with the Extracto CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this parser's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.Extracto.com/api/client/js/docs.md) (`npm install Extracto-client`) and [Python](https://docs.Extracto.com/api/client/python/docs.md) (`pip install Extracto-client`).

# README

## Shopify Scraper

[![Shopify products, every variant: prices, stock, barcodes, images, video, tags and subscriptions, with collections, keyword search, recommendations and multi-currency. Pay per result, not per minute.](https://api.Extracto.com/v2/key-value-stores/l7CL6aCEsAbA5My3a/records/readme-hero-v1.png)](https://console.Extracto.com/parsers/Us2dCgQWZ0A8L9prQ/input)

**Every product a Shopify store publishes, down to the variant.** Give it a store, a collection or a
product page, or a keyword to search the stores you list. Each product comes back with every variant's
price, stock status and SKU, and with its description, images, tags and options. Product pages, and
collection crawls with the matching option, add the merchant's barcode, weights and subscription plans.
With `grabRealInventory`, stores that publish their stock numbers give real counts. No API key needed.

> ๐Ÿงฐ Need the list of stores first? [Shopify Store Leads](https://extracto.cloud/docs/api/parsers/shopify) finds Shopify stores by country, category or keyword, ready for this parser to scrape.

### Why choose this parser?

**The whole catalog, not the first page.** Point it at a store's home page and it reads the entire catalog,
page after page, to the end. Point it at a collection or a single product and it reads just that. A keyword
search runs across every store you list. If a store has renamed a product, the scraper follows it to its
new page.

**More than Shopify's product JSON.** A product read from its own data carries more than prices, stock
status, images, tags and options. You also get the merchant's barcode (UPC, EAN or GTIN), weight, tax and
fulfilment settings, order quantity rules and B2B volume price breaks. Its storefront adds what no JSON
endpoint publishes: subscription and pre-order plans with their prices, and product videos. Every product
also gets a rolled-up `price`, a `priceRange` and a `fullyOutOfStock` flag. **Output** below says which
crawls return which fields.

**Real stock counts, where a store publishes them.** With `grabRealInventory`, each variant carries the
number in stock (`quantityAvailable`), not just in or out of stock. Many stores publish no stock numbers.
Those report the count as unknown and are never billed for it.

**Prices in the currency the store charged.** Prices are integers in minor units, so `2000` means 20.00.
`source.currency` labels them with the currency the store actually charged for the request. To get
another currency, add `?currency=` to the start URL, or use a Shopify Markets locale path such as `/nl-nl/`.

**Built for stores that push back.** A request that a store's bot check refuses is sent again from a fresh
proxy address. A rate-limited request waits as long as the store asks, up to a minute, and then tries
again. A site that is not a Shopify store fails at once, with a message that says so. A store or product
that fails never ends the run.

**Pay per result, not per minute.** A duplicate is never billed twice, and **Max result records** caps
how much a run saves. A run that reaches its **Max total charge** stops cleanly and keeps everything it
saved. **Pricing** below lists what each event costs.

**[Try it on your own store โ†’](https://console.Extracto.com/parsers/Us2dCgQWZ0A8L9prQ/input)**

### ๐Ÿš€ Quick start

1. Click **Try for free**. The input opens with one collection already filled in. Start it as it is to
   see what a record looks like.
2. Put the pages you want into **Shopify site start urls**: a store's home page for its whole catalog, a
   collection, or a single product. To check that a site runs on Shopify, open `https://<domain>/admin`.
   If you see a Shopify login page, the site is compatible.
3. Set **Max result records**. It caps how many records the run saves, and so the most the run can
   charge. The default is `100`; `0` means no limit.
4. Optionally, under **Extra data per product**, set **Max recommended products**, or turn on
   **Grab real inventory (experimental)** or **Grab storefront detail (selling plans & video)**.
5. Click **Start**, then download the results from the **Storage** tab as JSON, CSV, Excel or XML, or
   pull them from the API.

Start small, check that the records look the way you expect, then raise **Max result records**.

### ๐Ÿ’ก What people use it for

**Research a competitor's range.** Get a whole store with no cap on records: prices, variants, stock, and
how the store tags and categorises its products.

```json
{ "startUrls": [{ "url": "https://kith.com" }], "maxRequestsPerCrawl": 0 }
```

**Watch one category for launches and stock changes.** Crawl one collection on a schedule and compare
the runs:

```json
{ "startUrls": [{ "url": "https://kith.com/collections/kith-tops" }], "maxRequestsPerCrawl": 0 }
```

**Build one catalog across brands.** Crawl several stores in one run; each record's
`source.canonicalUrl` shows which store it came from:

```json
{ "startUrls": [{ "url": "https://kith.com" }, { "url": "https://www.allbirds.com" }], "maxRequestsPerCrawl": 0 }
```

**Check one product, variant by variant**: the price, stock and barcode of every size.

```json
{ "startUrls": [{ "url": "https://www.allbirds.com/products/mens-dasher-nz-anthracite" }] }
```

**See what a store recommends alongside a product:**

```json
{ "startUrls": [{ "url": "https://www.allbirds.com/products/mens-dasher-nz-anthracite" }], "maxRecommendationsPerProduct": 10 }
```

**Read real stock counts** from a store that publishes them:

```json
{ "startUrls": [{ "url": "https://www.allbirds.com/collections/mens" }], "grabRealInventory": true }
```

**Find products by keyword** in the stores you list:

```json
{ "startUrls": [{ "url": "https://kith.com" }], "query": "hoodie" }
```

**Collect prices in another currency.** On a store that does not use Shopify Markets, add `?currency=`:

```json
{ "startUrls": [{ "url": "https://kith.com/collections/kith-tops?currency=EUR" }] }
```

**For a store on Shopify Markets, use its local market instead**, with the locale in the path:

```json
{ "startUrls": [{ "url": "https://flyingtiger.com/nl-nl/collections/shop-all" }] }
```

### ๐Ÿ“‹ Input

The scraper accepts a JSON input defining the target URLs and scraping behavior.

| Parameter | Type | Required | Description |
| :--- | :--- | :--- | :--- |
| `startUrls` | Array | Yes | A list of URLs to scrape (Homepage, Collection, or Product URLs). |
| `proxy` | Object | No | Proxy configuration. Defaults to the Extracto proxy (datacenter). **Residential proxies are highly recommended** to avoid blocks. Pass `{ "useApifyProxy": false }` to scrape without one. |
| `maxRequestsPerCrawl` | Integer | No | Limit the number of records saved โ€” and therefore billed. Set to `0` for unlimited (default `100`). |
| `maxRecommendationsPerProduct` | Integer | No | Number of recommended products to fetch per product. Billed per recommended product โ€” see **Pricing** below. Set to `0` to disable (default `0`). Max `20`. |
| `grabRealInventory` | Boolean | No | Fetch real per-variant inventory counts from each product's JSON. Billed per product that returns a real count; adds one request per product on collection crawls. Only stores that expose inventory in their storefront JSON are supported โ€” others report it as unknown and are not billed for it (default `false`). |
| `maxInventoryFetches` | Integer | No | **Advanced, not shown in the form.** Cap on the extra per-product inventory fetches during collection crawls (only when `grabRealInventory` is enabled). Unlimited unless set; pass a number in the JSON input to cap the run. |
| `grabStorefrontDetail` | Boolean | No | On **collection crawls**, also fetch each product's storefront payload for selling plans, video media and the price span. Adds one request per product and is not billed as a separate event. Product URLs fetch this payload regardless โ€” it is where their stock availability comes from (default `false`). |
| `maxStorefrontFetches` | Integer | No | **Advanced, not shown in the form.** Cap on those extra per-product storefront fetches during collection crawls (only when `grabStorefrontDetail` is enabled). Unlimited unless set; pass a number in the JSON input to cap the run. |
| `query` | String | No | Search query to find specific products. Each start URL is searched in turn; a store that cannot be searched is skipped. |

### ๐Ÿ“ค Output

Data is stored in the default dataset. The scraper outputs either **Product** or **Collection** objects.

Every product record carries a rolled-up `price` (the cheapest variant), a `priceRange`
(`min`/`max`/`varies`) and a `fullyOutOfStock` flag, all derived from the variants themselves โ€” no extra
request, every store, every crawl path. Variants always carry `requiresShipping`, `taxable` and `weightGrams`.

`stockStatus` takes one of three values: `InStock`, `OutOfStock`, or `LowInStock`. The last is reported
only where a real stock count was read and it is 5 or fewer, so a store that publishes no stock numbers
never returns it.

A product URL whose product the store has since moved โ€” usually a rename โ€” is followed to its new page.
The record's `source.canonicalUrl` is that page and `source.redirectedFrom` holds the URL you gave. A store
can also point an old product address at a replacement product, so compare the two where it matters which
product you get.

#### Detail enrichment

When the scraper holds a product's own JSON, variants additionally carry:

| Field | What it is |
|---|---|
| `barcode` | The merchant's UPC, EAN or GTIN. Shopify stores all of them in this one field. |
| `weight`, `weightUnit` | The weight as the merchant entered it. `weightGrams` is the same value normalised. |
| `quantityRule` | Minimum, maximum and step size for ordering the variant. Shopify's default is `{min: 1, increment: 1}` with no maximum, and that is what the great majority of storefronts return โ€” read it as *no ordering rules configured*, not as a measured value. |
| `quantityPriceBreaks` | B2B volume pricing tiers โ€” buy *n* or more, pay this unit price. Empty on a typical consumer storefront. |
| `taxCode`, `fulfillmentService` | The merchant's tax code and who fulfils the variant. |

and the product carries `publishedScope` and `templateSuffix`. Image `alt` text is filled in here too, where the
store provides it.

**Which crawls get it.** A product URL always does โ€” the scraper already holds that product's JSON, so enrichment
costs no extra request and needs no option enabled, including with `grabRealInventory` off. A collection crawl
gets it whenever `grabRealInventory` is on, because the detail fetch that option pays for carries these fields
too, at no additional request and no additional charge. A plain collection crawl with `grabRealInventory` off
fetches no product JSON and so returns none of them.

#### Storefront detail

A product's storefront payload publishes things none of the JSON endpoints do:

| Field | What it is |
|---|---|
| `sellingPlanGroups`, `requiresSellingPlan` | Subscriptions, pre-orders and deposit plans the merchant offers, with each plan's options. `requiresSellingPlan` is true when the product cannot be bought outright. |
| `sellingPlanAllocations` (per variant) | What that variant costs under each plan, including the per-delivery price on a recurring one. |
| Video media | `medias` entries typed `Video`. The other endpoints publish images only, so without this option a product's videos are invisible. |

**Which crawls get it.** A product URL always does, with no option to enable: that payload is where a
product URL's stock availability comes from, because the single-product JSON carries no availability flag
at all. A collection crawl gets it when `grabStorefrontDetail` is on โ€” there the availability is already in
the listing, so the payload is bought purely for the fields above. It costs one extra request per product,
is not billed as a separate event, and is off by default; `maxStorefrontFetches` caps it, since one start
URL can mean hundreds of products.

**How availability is resolved on a product URL.** Three sources, strongest first: the storefront payload's
per-variant flag (joined on variant id), then the product page's embedded schema.org data (joined on SKU),
then the JSON API's own flag. The page is only fetched when the payload yields nothing for any variant โ€” a
partial answer is kept as-is rather than paying for a rendered page, which is roughly seventy times more
data than the payload.

> **Inventory is narrower than enrichment.** `quantityAvailable` (real stock count), `inventoryTracked` and
> `inventoryPolicy` need `grabRealInventory` *and* a store that publishes stock numbers in its storefront JSON โ€”
> many do not, and those are simply omitted and not billed. The enrichment fields above come from the same
> request but are published by a wider set of stores, so a store that reveals no stock counts at all can still
> return complete barcode, weight and quantity-rule data.

<details>
<summary><strong>View Product Output Example</strong></summary>

```json
{
 "source": {
  "id": "7205191974992",
  "canonicalUrl": "https://www.allbirds.com/products/mens-dasher-nz-anthracite",
  "retailer": "Allbirds",
  "language": "en",
  "currency": "USD",
  "createdUTC": 1754531673000,
  "updatedUTC": 1789383693000,
  "publishedUTC": 1787590381000
 },
 "title": "Men's Dasher NZ - Anthracite (Dark Anthracite Sole)",
 "description": "<p>A new take on our fan-favorite Dasher, made for busy days and spontaneous plans. Lightweight, breathable comfort keeps you cool as you move, with added heel protection where it counts. Fast when you want. Comfortable always.</p>",
 "brand": "Allbirds",
 "categories": [
  "Shoes"
 ],
 "tags": [
  "allbirds::carbon-score => undefined",
  "allbirds::cfId => color-ec0dcc8441508c674c885b8c8a990cfa",
  "allbirds::complete => true",
  "allbirds::edition => limited",
  "allbirds::gender => mens",
  "allbirds::hue => grey",
  "allbirds::master => mens-dasher-nz",
  "allbirds::material => tree",
  "allbirds::price-tier => msrp",
  "allbirds::silhouette => dasher",
  "FLEX3818571 (PL 553696876312)",
  "loop::returnable => true",
  "OOS DNS",
  "WAVE 2 WCDC",
  "YCRF_mens-perform-shoes",
  "YGroup_ygroup_mens-dasher-nz"
 ],
 "variants": [
  {
   "id": "41271176757328",
   "title": "8",
   "sku": "A12417M080",
   "options": [
    "8"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "OutOfStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 914,
   "barcode": "196942279229",
   "weight": 2.0154,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 0
  },
  {
   "id": "41271176790096",
   "title": "8.5",
   "sku": "A12417M085",
   "options": [
    "8.5"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "OutOfStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 927,
   "barcode": "196942279281",
   "weight": 2.044,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 0
  },
  {
   "id": "41271176822864",
   "title": "9",
   "sku": "A12417M090",
   "options": [
    "9"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 936,
   "barcode": "196942279373",
   "weight": 2.0639,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 17
  },
  {
   "id": "41271176855632",
   "title": "9.5",
   "sku": "A12417M095",
   "options": [
    "9.5"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 962,
   "barcode": "196942279410",
   "weight": 2.1212,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 48
  },
  {
   "id": "41271176888400",
   "title": "10",
   "sku": "A12417M100",
   "options": [
    "10"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 978,
   "barcode": "196942279540",
   "weight": 2.1565,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 68
  },
  {
   "id": "41271176921168",
   "title": "10.5",
   "sku": "A12417M105",
   "options": [
    "10.5"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1005,
   "barcode": "196942279670",
   "weight": 2.216,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 80
  },
  {
   "id": "41271176953936",
   "title": "11",
   "sku": "A12417M110",
   "options": [
    "11"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1044,
   "barcode": "196942279786",
   "weight": 2.302,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 40
  },
  {
   "id": "41271176986704",
   "title": "11.5",
   "sku": "A12417M115",
   "options": [
    "11.5"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1056,
   "barcode": "196942279892",
   "weight": 2.3285,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 19
  },
  {
   "id": "41271177019472",
   "title": "12",
   "sku": "A12417M120",
   "options": [
    "12"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1081,
   "barcode": "196942280003",
   "weight": 2.3836,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 17
  },
  {
   "id": "41271177052240",
   "title": "12.5",
   "sku": "A12417M125",
   "options": [
    "12.5"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "OutOfStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1087,
   "barcode": "196942280096",
   "weight": 2.3968,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 0
  },
  {
   "id": "41271177085008",
   "title": "13",
   "sku": "A12417M130",
   "options": [
    "13"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "OutOfStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1141,
   "barcode": "196942280164",
   "weight": 2.5159,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 0
  },
  {
   "id": "41271177150544",
   "title": "14",
   "sku": "A12417M140",
   "options": [
    "14"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "InStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1199,
   "barcode": "196942280317",
   "weight": 2.6438,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 7
  },
  {
   "id": "41271177183312",
   "title": "15",
   "sku": "A12417M150",
   "options": [
    "15"
   ],
   "price": {
    "current": 14000,
    "previous": 0,
    "stockStatus": "OutOfStock"
   },
   "requiresShipping": true,
   "taxable": true,
   "weightGrams": 1199,
   "barcode": "196942280393",
   "weight": 2.6438,
   "weightUnit": "lb",
   "taxCode": "PC040144",
   "fulfillmentService": "manual",
   "quantityRule": {
    "min": 1,
    "increment": 1
   },
   "inventoryTracked": true,
   "inventoryPolicy": "deny",
   "quantityAvailable": 0
  }
 ],
 "medias": [
  {
   "id": "35670548119632",
   "type": "Image",
   "url": "https://cdn.shopify.com/s/files/1/1104/4168/files/A12416_26Q1_Dasher-NZ-Anthracite-Dark-Anthr_PDP_LEFT.png?v=1768948005",
   "variantIds": [],
   "alt": ""
  },
  {
   "id": "35670548152400",
   "type": "Image",
   "url": "https://cdn.shopify.com/s/files/1/1104/4168/files/A12416_26Q1_Dasher-NZ-Anthracite-Dark-Anthr_PDP_BACK.png?v=1768948005",
   "variantIds": [],
   "alt": ""
  },
  {
   "id": "35670548185168",
   "type": "Image",
   "url": "https://cdn.shopify.com/s/files/1/1104/4168/files/A12416_26Q1_Dasher-NZ-Anthracite-Dark-Anthr_PDP_TD.png?v=1768948006",
   "variantIds": [],
   "alt": ""
  },
  {
   "id": "35670548217936",
   "type": "Image",
   "url": "https://cdn.shopify.com/s/files/1/1104/4168/files/A12416_26Q1_Dasher-NZ-Anthracite-Dark-Anthr_PDP_SOLE.png?v=1768948006",
   "variantIds": [],
   "alt": ""
  },
  {
   "id": "35670548250704",
   "type": "Image",
   "url": "https://cdn.shopify.com/s/files/1/1104/4168/files/A12416_26Q1_Dasher-NZ-Anthracite-Dark-Anthr_PDP_PAIR_3Q.png?v=1768948006",
   "variantIds": [],
   "alt": ""
  }
 ],
 "options": [
  {
   "type": "Size",
   "values": [
    {
     "id": "8",
     "name": "8"
    },
    {
     "id": "8.5",
     "name": "8.5"
    },
    {
     "id": "9",
     "name": "9"
    },
    {
     "id": "9.5",
     "name": "9.5"
    },
    {
     "id": "10",
     "name": "10"
    },
    {
     "id": "10.5",
     "name": "10.5"
    },
    {
     "id": "11",
     "name": "11"
    },
    {
     "id": "11.5",
     "name": "11.5"
    },
    {
     "id": "12",
     "name": "12"
    },
    {
     "id": "12.5",
     "name": "12.5"
    },
    {
     "id": "13",
     "name": "13"
    },
    {
     "id": "14",
     "name": "14"
    },
    {
     "id": "15",
     "name": "15"
    }
   ]
  }
 ],
 "price": {
  "current": 14000,
  "previous": 0,
  "stockStatus": "InStock"
 },
 "priceRange": {
  "min": 14000,
  "max": 14000,
  "varies": false
 },
 "fullyOutOfStock": false,
 "publishedScope": "global",
 "templateSuffix": "mens-dasher-nz",
 "requiresSellingPlan": false
}
```

</details>

<details>
<summary><strong>View Collection Output Example</strong></summary>

```json
{
 "id": "448305135744",
 "title": "&Kin Winter 2025",
 "handle": "kin-winter-2025",
 "description": "<p>An exclusive partnership with Giorgio Armani as the official suiting partner for &amp;Kin, reintroducing the Traveller and the Artist silhouettes. Discover unbranded staples made from ultra-soft cashmere, pony hair leather, extra-fine merino wool, and luxe French terry, as well as textured finishes such as Basketweave knit and Tile Herringbone.ย <br></p>",
 "productsCount": 30,
 "updatedUTC": 1766923230000,
 "publishedUTC": 1761312567000
}
```

</details>

### ๐Ÿ’ณ Pricing

This parser uses Extracto's **pay-per-event** model: you are charged for the results it produces, not for how long it runs.

| Event | Price | Charged | Notes |
| :--- | ---: | :--- | :--- |
| `product` | $0.0018 | Once per product saved to the dataset | The main unit of value. Volume tiers below. |
| `collection` | $0.0008 | Once per collection saved to the dataset | Only `/collections` crawls emit these. |
| `recommends` | $0.0010 | Once per recommended product | Off by default (`maxRecommendationsPerProduct` is `0`). |
| `real-inventory` | $0.0012 | Once per product with a real stock count | Off by default (`grabRealInventory`). |

`product` is tiered by your total monthly Extracto spend, so the more you run the less each product costs:

| Tier | FREE | BRONZE | SILVER | GOLD | PLATINUM | DIAMOND |
| :--- | ---: | ---: | ---: | ---: | ---: | ---: |
| Per product | $0.0018 | $0.0016 | $0.0014 | $0.0012 | $0.0012 | $0.0012 |

At the GOLD tier, 100,000 products costs $120.

<details>
<summary>What each event means in detail</summary>

| Event | Charged | Notes |
| :--- | :--- | :--- |
| `product` | Once per product saved to the dataset | The main unit of value. |
| `collection` | Once per collection saved to the dataset | Only `/collections` crawls emit these โ€” they list a store's categories rather than its products, and a single request can return up to 250 of them. |
| `recommends` | Once per recommended product | Recommendations are nested inside the parent product record rather than saved as separate results, so they are billed here. Only applies when `maxRecommendationsPerProduct` is above `0`. |
| `real-inventory` | Once per product that comes back with a real stock count | Only applies when `grabRealInventory` is enabled **and** the store actually exposes inventory. Stores that don't expose it are never billed for it. |

</details>

#### Controlling your spend

- **`maxRequestsPerCrawl`** caps how many records are saved, so it caps `product`/`collection` charges directly. Default `100`; `0` means unlimited.
- **Max total charge** (set per run or per task in the Extracto UI) is a hard ceiling. When it is reached the parser stops of its own accord and the run finishes with everything it had already saved โ€” it is not cut off mid-fetch.
- **Duplicates are never billed twice.** A product reached through several start URLs or through overlapping collections, or re-visited after a failed request is retried, is saved and charged once per run.

### โš ๏ธ What this parser does not do

- **It does not merge SKUs.** On certain Shopify websites, multiple SKUs may be visually merged into a single product page. This scraper treats each SKU as an individual product entry and does not perform any merging of these SKUs.
- **It does not return 3D or AR models.** Shopify also publishes 3D/AR model media. The output spec has no type for it, so those entries are skipped rather than mislabelled as images.
- **It does not guess stock counts.** A real count needs `grabRealInventory` and a store that publishes
  its stock numbers. A store that does not publish them gets no count and is not billed for one.

### ๐Ÿงฐ Other parsers by autofacts

Extracto only auto-recommends parsers in the same category, so here are the ones that actually pair with this scraper:

| parser | What it's for |
| :--- | :--- |
| [Shopify Store Leads](https://extracto.cloud/docs/api/parsers/shopify) | Find and qualify the stores first โ€” catalog size, apps, theme, contacts โ€” then feed the domains into this parser |
| [Schema Markup Scraper & SEO Auditor](https://extracto.cloud/docs/api/parsers/shopify) | Audit a store's structured data, Open Graph tags and canonical setup |
| [Google Ads Scraper](https://extracto.cloud/docs/api/parsers/shopify) | The Google ads a store runs โ€” search, Shopping and YouTube โ€” with the landing pages they lead to |
| [Sephora Product Scraper](https://extracto.cloud/docs/api/parsers/shopify) | Non-Shopify beauty retail, 29 countries |
| [Macy's Scraper](https://extracto.cloud/docs/api/parsers/shopify) | Department-store catalog and pricing |
| [Universal Web Printer](https://extracto.cloud/docs/api/parsers/shopify) | Render any product page to PDF/PNG for archiving or evidence |
| [WooCommerce Scraper](https://extracto.cloud/docs/api/parsers/shopify) | The same job on WooCommerce stores, with a matching record shape |

All of them: [Extracto.com/autofacts](https://Extracto.com/autofacts)

***

### ๐Ÿ› ๏ธ Troubleshooting

| Issue | Possible Cause | Solution |
| :--- | :--- | :--- |
| **0 Results Found** | The site may not be Shopify-based or has strong anti-bot protection. | Verify with the `/admin` trick. Try using residential proxies. |
| **Access Denied / 403** | Your IP has been flagged. | Enable `useApifyProxy` and ensure you have sufficient proxy quota. |
| **Incorrect Price** | Raw integer format. | Prices are integer minor units โ€” 2000 means 20.00 in `source.currency`. |
| **Wrong currency** | The store runs Shopify Markets, where `?currency=XYZ` is ignored. | Use the store's locale-prefixed URL instead (e.g. `/nl-nl/collections/...`), which selects the market and its currency. Verified on one Markets store; `?currency=` still works on stores that are not using Markets. |
| **Run stopped early** | The run's **Max total charge** limit was reached. | Expected behaviour โ€” the parser stops cleanly and keeps everything it saved. Raise the limit, or lower `maxRequestsPerCrawl` to fit the budget. |

***

### ๐Ÿค– Use with AI agents

This parser is callable as a tool by any MCP-capable agent โ€” Claude, Cursor, VS Code โ€” or by your
own code, with no wrapper and nothing extra to deploy.

**Connect over MCP**

```
https://mcp.Extracto.com?tools=autofacts/shopify
```

In a client that reads an `mcpServers` configuration block:

```json
{
  "mcpServers": {
    "Extracto": {
      "url": "https://mcp.Extracto.com?tools=autofacts/shopify",
      "headers": { "Authorization": "Bearer YOUR_APIFY_TOKEN" }
    }
  }
}
```

The agent reads this parser's parameters and their descriptions straight from the input
schema, and the hosted server infers the result field types from the dataset schema โ€” so a
model knows what to send and what comes back before it ever calls anything.

**Or call the API directly**

```bash
curl -X POST "https://extracto.cloud/api/v1/parsers/shopify" \
  -H 'Content-Type: application/json' \
  -d '{"startUrls": [{"url": "https://kith.com/collections/kith-tops"}]}'
```

The response body is the dataset records described above.

# Changelog

This parser's version history is a separate document: https://extracto.cloud/docs/api/parsers/shopify/changelog.md

# parser input Schema

## `startUrls` (type: `array`):

Start URLs of Shopify site to start the parser. Site root url, category page urls, product page urls are all supported.

## `query` (type: `string`):

Search query to find products.

## `maxRequestsPerCrawl` (type: `integer`):

The maximum number of records saved to the dataset. Each record is one billed result, so this is also the run's cost cap. The scraper stops when the limit is reached. <br><br>If set to <code>0</code>, there is no limit โ€” use the run's <b>Max total charge</b> setting to bound spend instead.

## `maxResults` (type: `integer`):

Maximum number of results to return. Hidden parameter, overrides maxRequestsPerCrawl if set.

## `maxRecommendationsPerProduct` (type: `integer`):

The maximum number of recommended products to fetch per product. Recommendations are returned nested inside the product record and are billed per recommended product, so a value of 20 can cost up to 20 extra events per product. Set to 0 to disable. Default is 0, max is 20.

## `grabRealInventory` (type: `boolean`):

Fetch each product's real per-variant inventory count from its product JSON. Billed as a separate event for every product that comes back with a real stock count. On collection crawls it also adds one request per product. Only works for stores that expose inventory in their storefront JSON โ€” others report inventory as unknown, and are not billed for it. Off by default.

## `maxInventoryFetches` (type: `integer`):

Hard cap on the extra per-product detail fetches used to read inventory on collection crawls. Advanced, not shown in the form: unlimited unless set. Pass a number in the JSON input to cap the run. Only applies when 'Grab real inventory' is enabled.

## `grabStorefrontDetail` (type: `boolean`):

Fetch each product's storefront payload for selling plans (subscriptions, pre-orders), video media, and the product's price span. <b>Applies to collection crawls only</b> โ€” product URLs always fetch this payload, because it is where their stock availability comes from. Costs one extra request per product on collection crawls, and is not billed as a separate event. Off by default.

## `maxStorefrontFetches` (type: `integer`):

Hard cap on the extra per-product storefront fetches on collection crawls. Advanced, not shown in the form: unlimited unless set. Pass a number in the JSON input to cap the run. Only applies when 'Grab storefront detail' is enabled; product URLs are already bounded by the record limit.

## `grabContentSections` (type: `boolean`):

Also return the labelled information panels the store shows outside the description - tabs and accordions such as Features, Includes, Materials or Care - as the site renders them, in a `contentSections` field. Labels are kept exactly as shown, so panels like Shipping or Returns come along too and are yours to filter. Costs one extra page request per product on stores that have them; stores that have none are detected and skipped. Billed per product that actually returns panels.

## `maxContentSectionFetches` (type: `integer`):

Caps how many product pages the panel extraction may fetch in one run. Advanced, not shown in the form: unlimited unless set. Pass a number in the JSON input to cap the run. Only applies when "Grab product information panels" is on; a product URL is not capped by this.

## `proxy` (type: `object`):

Select proxies to be used by your crawler.

## parser input object example

```json
{
  "startUrls": [
    {
      "url": "https://kith.com/collections/kith-tops"
    }
  ],
  "maxRequestsPerCrawl": 100,
  "maxResults": 0,
  "maxRecommendationsPerProduct": 0,
  "grabRealInventory": false,
  "maxInventoryFetches": 0,
  "grabStorefrontDetail": false,
  "maxStorefrontFetches": 0,
  "grabContentSections": false,
  "maxContentSectionFetches": 0,
  "proxy": {
    "useApifyProxy": true
  }
}
```

# parser output Schema

## `products` (type: `string`):

No description

# API

You can run this parser programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'Extracto-client';

// Initialize the ApifyClient with your Extracto API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare parser input
const input = {
    "startUrls": [
        {
            "url": "https://kith.com/collections/kith-tops"
        }
    ],
    "proxy": {
        "useApifyProxy": true
    }
};

// Run the parser and wait for it to finish
const run = await client.parser("autofacts/shopify").call(input);

// Fetch and print parser results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`๐Ÿ’พ Check your data here: https://console.Extracto.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// ๐Ÿ“š Want to learn more ๐Ÿ“–? Go to โ†’ https://docs.Extracto.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Extracto API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the parser input
run_input = {
    "startUrls": [{ "url": "https://kith.com/collections/kith-tops" }],
    "proxy": { "useApifyProxy": True },
}

# Run the parser and wait for it to finish
run = client.parser("autofacts/shopify").call(run_input=run_input)

# Fetch and print parser results from the run's dataset (if there are any)
print(f"๐Ÿ’พ Check your data here: https://console.Extracto.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# ๐Ÿ“š Want to learn more ๐Ÿ“–? Go to โ†’ https://docs.Extracto.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://kith.com/collections/kith-tops"
    }
  ],
  "proxy": {
    "useApifyProxy": true
  }
}' |
Extracto call autofacts/shopify --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "Extracto": {
            "type": "http",
            "url": "https://mcp.Extracto.com/?tools=fetch-parser-details,autofacts/shopify"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Extracto Console (https://console.Extracto.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.Extracto.com/v2/parsers/Us2dCgQWZ0A8L9prQ/builds/135K2vebS4RIGiRx3/openapi.json

Request Body

{
  "input": {
    "startUrls": [
        {
            "url": "https://kith.com/collections/kith-tops"
        }
    ]
}
}

cURL

curl -X POST https://extracto.cloud/api/v1/parsers/shopify \
  -H "X-API-KEY: ext_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{"input":{"startUrls":[{"url":"https://kith.com/collections/kith-tops"}]}}'