Collector

Shopify Storefronts scraper

Every product a Shopify store publishes — title, price, compare-at price, stock, SKU, vendor, type, tags and images, a row per product, straight off the storefront’s own catalogue file with no account and no browser.

On this page

What it collects

commerce

Every product a Shopify store publishes — title, price, compare-at price, stock, SKU, vendor, type, tags and images, a row per product, straight off the storefront’s own catalogue file with no account and no browser. It is a scraper, which means the whole contract is: fill in the form, and one table comes back. No browser window opens, it costs none of your plan's automation slots, and the requests go out over plain HTTP carrying a real browser's TLS fingerprint.

It runs fine from your own address; a proxy is available on the run and this target does not need one. Whichever way it went, every row records it — see the provenance columns below.

What you give it — 7

The form the app draws, read out of the card itself. Each hint below is the one shown beside that field in the run dialog, so nothing here is a description of the product written separately from it.

The one answer it needs

startUrls
Stores · list · required

One store per line — the domain, the homepage, or any product or collection link from it. A 404 means that host does not serve this file: either it is not a Shopify store, or it is one that has switched the catalogue off. The collector is fine either way and the other stores keep going.

Everything else is optional

maxProductsPerStore
Products per store · number

Rounded up to whole pages of 250. There is no early stop — a shop with forty products still costs the page count asked for, so raise this for a big catalogue rather than leaving it high for everything.

keywords
Title contains · list

Keeps a product only when its title contains one of these. Costs no extra requests — the reply already carried the title.

productTypes
Product type · list

Matched against the shop’s own type, exactly as it writes it.

vendors
Vendor · list

Useful on a reseller, where one storefront carries many brands.

minPrice
Minimum price · number

In the store’s own currency, which this endpoint does not name. A product with no price fails this rather than counting as zero.

onlyAvailable
In stock only · boolean

Drops products whose first variant is sold out.

What comes back — 22 columns

One dataset per run — a typed table your workspace owns, which you can then sort, filter, edit in the grid, export whole, or read back over the local API. These are the columns it is created with.

3 columns are marked may be empty. That is a fact about the record rather than about the collector — a post with nowhere tagged, an account with no business category — and it is said out loud so a blank cell does not read as a broken scraper. A column the reply never carries at all is deleted upstream rather than shipped empty, so nothing here is decoration.

18 from Shopify Storefronts

KeyColumnType
titleProducttext
productUrlProduct pageurl
vendorBrandtext
productTypeTypemay be emptytext
pricePricenumber
compareAtPriceCompare-at pricemay be emptynumber
availableIn stockselect
skuSKUtext
variantCountVariantsnumber
tagsTagsmay be emptytext
imageUrlImageurl
imageCountImagesnumber
handleHandletext
publishedAtPublisheddatetime
createdAtCreateddatetime
updatedAtUpdateddatetime
productIdProduct IDtext
searchTermStoretext

4 that every scraper writes

The same four on every card, so a table can still answer months later how its rows got there: which service they came from, when, which profile's identity the requests carried, and whether that identity was signed in.

KeyColumnType
platformPlatformtext
collected_atCollected atdatetime
profileCollected byprofile
logged_inSigned inselect

Four ways to run it

The Scrapers tab. Pick the card, fill in the form, press Start. Every run started here makes a new table, named after what you searched for and when. Check first if you like — it collects one page, writes nothing, and reports which columns came back filled.

Ask the assistant. It has the whole catalogue, so this card is a sentence rather than a form. It fills in the parameters above from what you said and shows you them before it starts.

Over MCP or the local API. The same card, from a coding agent — argus_run_scraper over the MCP server, or POST http://127.0.0.1:39219/v1/scrapers/run on the local API. There is a sample call beside it that writes nothing.

argus_run_scraper — Shopify Storefronts
{
  "kind": "shopify_products",
  "inputs": {
    "startUrls": [
      "…"
    ]
  }
}

As a step inside a workflow. The Run scraper step puts this collector in the middle of an automation — collect, then filter, then mail — and the tree still opens no window doing it. It is also the one caller that can overrule the table-per-run rule: pointed at a table you name, it creates that table with the columns above if it is missing and files every run into it after, appending or updating the row in place on a column you match on. The step hands the next one the table's name, its id and the row count.

Where it stops

It will not drive a page. Anything that needs a real browser — a site with no list endpoint behind it, or anything operating an account of your own — is an automation instead, and that is a different tool rather than a worse version of this one.

It signs in to nothing. Reading a public page signed out is the settled position — Bright Data scraped Meta signed out and won; hiQ scraped LinkedIn signed in and lost — so this card asks for no account and holds none.