采集卡片

Indeed Jobs 采集器

在 Indeed 上针对某个搜索投放的每一家雇主:岗位、公司、公司的评分和评价数、地点、薪资区间、工作类型和发布多久了,每个职位一行,不需要账号,也不开浏览器。

本页内容

它采什么

招聘建议挂代理

在 Indeed 上针对某个搜索投放的每一家雇主:岗位、公司、公司的评分和评价数、地点、薪资区间、工作类型和发布多久了,每个职位一行,不需要账号,也不开浏览器。 它是一张采集器卡片,也就是说全部约定只有一条:填好表单,回来一张表。不会打开浏览器窗口,不占用你套餐里的任何自动化名额,请求走的是普通 HTTP,带着真实浏览器的 TLS 指纹。

量一大,这个目标就会在意请求是从哪里发出的,所以运行时请选一个代理——这是表单上的建议而不是门槛,不选也照样能跑。无论走的是哪条路,每一行都会记下来——见下面的溯源列。

你要提供什么——10 项

应用画出的那张表单,直接读自这张卡片本身。下面每一条提示,都是运行对话框里显示在该字段旁边的那一条,所以这里没有任何脱离产品另写一遍的说明。

它唯一需要的那个答案

searches
Searches · list · 必填

One search per line, exactly what you would type into Indeed. Each costs its own walk, and the rows are pooled and de-duplicated on the job key, so two overlapping searches do not double a job.

其余的都可填可不填

location
Where · text

A town, a city or a postcode, as Indeed’s own "where" box takes it. Leave it empty to search the whole country.

country
Country · select

Indeed runs a separate site per country and each searches its own market. This picks the site, and with it the language the listings come back in.

resultsLimit
Results per search · number

Rounded up to whole pages of 10. The walk stops early on its own when a page brings nothing new, so a search with forty results costs four requests and not the number asked for.

minSalary
Minimum salary · number

Matched against the top of the advertised range, in the site’s own currency. Most listings quote no salary at all and fail this — a job with no pay quoted is not a job paying zero.

minRating
Minimum company rating · number

Out of five, from Indeed’s own employee reviews.

companies
Company contains · list

Keeps a row only when the employer’s name contains one of these.

jobTypes
Job type contains · list

Matched against Indeed’s own words. Most listings do not state a type, and this filter drops them.

hideExpired
Hide closed listings · boolean

Indeed keeps returning some listings after they close.

onlyUrgent
Urgently hiring only · boolean

Keeps only the roles Indeed badges as urgent.

会拿回什么——27 列

每次运行生成一个数据集——一张带类型的表,归你的工作区所有:可以排序、筛选、在网格里直接编辑、整张导出,也可以通过本地 API 读回来。下面就是它创建时带的列。

有 6 列标着可能为空。那是关于记录本身的事实,而不是关于采集器的——一条没有标地点的帖子、一个没有商业类目的账号——把它说出来,是为了让空单元格不被读成采集器坏了。至于返回里根本不带的列,在上游就被删掉了,不会空着交付,所以这里没有一列是装饰。

来自 Indeed Jobs 的 23 列

键名类型
titleRoletext
companyCompanytext
companyRatingRatingnumber
companyReviewCountReviewsnumber
locationLocationtext
cityCitytext
regionRegiontext
postcodePostcode可能为空text
remoteWorkRemote可能为空text
jobTypesJob type可能为空text
salaryMinSalary from可能为空number
salaryMaxSalary to可能为空number
salaryPeriodSalary period可能为空text
postedPostedtext
easyApplyEasy applyselect
urgentlyHiringUrgentselect
sponsoredSponsoredselect
featuredEmployerFeaturedselect
expiredExpiredselect
jobUrlJob pageurl
companyPageCompany on Indeedurl
jobkeyJob keytext
searchTermSearchtext

每个采集器都会写的 4 列

每张卡片上都是同样这四列,好让一张表在几个月后仍能回答它的行是怎么来的:来自哪个服务、什么时候、请求带的是哪个环境的身份,以及那个身份当时是否已登录。

键名类型
platformPlatformtext
collected_atCollected atdatetime
profileCollected byprofile
logged_inSigned inselect

跑它的四种方式

「采集器」标签页。 挑中这张卡片,填好表单,按「开始」。从这里发起的每一次运行都会生成一张新表,以你搜索的内容和时间命名。想先试试也可以——它只采一页,什么都不写,并告诉你哪些列拿回了内容。

问助手。 它手里有整个目录,所以这张卡片是一句话,而不是一张表单。它会按你说的话把上面那些参数填好,并在开跑之前拿给你看。

通过 MCP 或本地 API。 同一张卡片,从编程 Agent 里调用——MCP 服务器上的 argus_run_scraper,或者本地 APIPOST http://127.0.0.1:39219/v1/scrapers/run。旁边有一个什么都不写入的示例调用。

argus_run_scraper — Indeed Jobs
{
  "kind": "indeed_jobs",
  "inputs": {
    "searches": [
      "…"
    ]
  }
}

作为工作流里的一个步骤。 Run scraper 步骤把这个采集器放进一条自动化流程的中间——先采、再筛、再发信——整棵流程照样一个窗口都不开。它也是唯一能推翻「一次运行一张表」规则的调用方:指定一张你命名的表,它会在表不存在时按上面的列建好,之后每次运行都写进这张表,追加或者按你指定的匹配列就地更新。这个步骤会把表名、表 id 和行数交给下一步。

它到哪里为止

它不会去驱动页面。凡是需要真实浏览器的事——背后没有列表接口的网站,或者任何操作你自己账号的事——都该交给自动化流程。那是另一件工具,而不是这件工具的劣化版。

它不登录任何地方。未登录读取公开页面是已成定论的做法——Bright Data 未登录采集 Meta 并且胜诉,hiQ 登录后采集 LinkedIn 则败诉——所以这张卡片不要账号,也不持有账号。

其他采同类东西的卡片

服务不同,最后拿到的表却是同一个形状。