Collect Crunchbase Builder data and return results
数据分析
"dataify-linkedin-company-information-by-url"
试用Collect Linkedin Builder data and return results
它能做什么
Collect structured LinkedIn company information from one or more known company URLs. Do not use for personal profiles, jobs, or Crunchbase URLs.
技能文档
Dataify Builder Skill
Use this skill to prepare Dataify builder requests for the scraper family rooted at linkedin_company_information_by-url on linkedin.com.
Workflow
- Check whether
DATAIFY_API_TOKENexists in the environment. - If the token is missing, stop and tell the user to sign in at Dataify Dashboard to obtain it.
- Ask the user to choose exactly one tool from the following Chinese list:
- 通过URL采集 (linkedin_company_information_by-url)
- 通过职位列表URL采集 (linkedin_job_listings_information_by-job-listing-url)
- 通过职位URL采集 (linkedin_job_listings_information_by-job-url)
- 通过关键词采集 (linkedin_job_listings_information_by-keyword)
- Read
references/tool-params.jsonand find the chosen tool bytool_signor Chinese tool name. - For each parameter in the chosen tool:
- If
input_modeisuser_input, ask the user for the value. - If
input_modeisselect, present the saved options to the user.
- If
- Use
scripts/build-dataify-request.pyas the default cross-platform helper. - Use
scripts/build-dataify-request.ps1as the Windows PowerShell helper when needed. - When a selectable parameter has a human-readable Chinese label, keep that label in
spider_parameters. Do not replace it with a code such asHKunless the user explicitly asks for the coded value. - Build
spider_parametersas a JSON array. - If every parameter has only one final value, build one object such as
[{"searchurl":"...","country":"Hong Kong"}]. - If one or more parameters have multiple aligned values, zip them by index and build one object per row. Example:
[{"search_url":"url1","page_turning":"1","max_num":"15"},{"search_url":"url2","page_turning":"1","max_num":"15"}]. - If a parameter has one value while another parameter has multiple values, reuse the single value across every generated row.
- Set
spider_nametolinkedin.com. - Set
spider_idto the selected tool'stool_sign. - Always include
spider_errors=trueandfile_name={{TasksID}}. - Return a curl command for
https://scraperapi.dataify.com/builder.
Set DATAIFY_API_TOKEN
Prefer a permanent environment-variable setup instead of setting the token only for the current terminal session.
Windows PowerShell, permanent for the current user:
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")
Then reopen PowerShell. If the current session also needs the token immediately, run:
$env:DATAIFY_API_TOKEN = "your_token_here"
macOS or Linux, permanent for bash:
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc
macOS or Linux, permanent for zsh:
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc
Script usage
Python:
python scripts/build-dataify-request.py --tool-sign --values-file values.json
PowerShell:
& ".\scripts\build-dataify-request.ps1" -ToolSign "" -ValuesFile ".\values.json"
The values.json file should contain either one object or an array of objects. Example:
[{"searchurl":"https://www.airbnb.com/s/Greece/homes?...","country":"Hong Kong"}]
Required output shape
Generate a curl command in this form:
curl -X POST 'https://scraperapi.dataify.com/builder' \
-H "Authorization: Bearer $DATAIFY_API_TOKEN" \
-H 'Content-Type: application/x-www-form-urlencoded' \
-d 'spider_name=linkedin.com' \
-d 'spider_id=' \
-d 'spider_parameters=[{"param":"value"}]' \
-d 'spider_errors=true' \
-d 'file_name={{TasksID}}'
Reference usage
references/tool-params.jsonstores the full saved parameter catalog for every available tool in this scraper family.scripts/build-dataify-request.pyis the portable implementation and should be preferred.scripts/build-dataify-request.ps1mirrors the same behavior for Windows users.- If a parameter has no options, the user must provide the value.
- If a parameter has options, present those options back to the user before building the final request.
- Do not assume
spider_parametersalways contains exactly one object. Multi-value tools may require multiple objects zipped by index. - Use the saved
url_exampleonly as a reference example. Do not assume the user wants the example values unless they explicitly confirm them.
相关技能
Looks up LinkedIn company, product, and showcase pages by ID via the Crawlora API, returning clean JSON. Use when the user wants a company's LinkedIn profile info, a product page, or a showcase page — instead of scraping LinkedIn directly. Covers company/product/showcase pages only, not personal LinkedIn profiles.
Collect Glassdoor Builder data and return results
Pull LinkedIn person and company profiles, their posts, job listings and post comments as structured JSON. 9 endpoints from 1 to 30 credits, four of them paginated, for prospecting, recruiting, and market research.
Analyze LinkedIn users and companies through the KeyAPI REST API using live official docs. Use for professional profiles, contact information, work experience, education, skills, publications, certifications, honors, recommendations, interests, posts, comments, videos, company profiles, employees, j
LinkedIn (linkedin.com). Use this skill for ANY LinkedIn request — reading, creating, updating, and deleting data. Whenever a task involves LinkedIn, use thi...