Skip to main content
POST
Crawl a website
1 credit per page Rate limit weight: 10 See the guide for examples and usage.
Need to crawl more than 500 pages? Use batches.

Authorizations

Authorization
string
header
required

Send Authorization: Bearer <API_KEY>. Keys have full access unless restricted to scopes.

Body

application/json
url
string<uri>
required

Start URL, including http:// or https://.

maxPages
integer
default:100

Maximum pages to crawl.

Required range: 1 <= x <= 500
maxDepth
integer

Maximum link depth from the starting URL (0 = only the starting page)

Required range: x >= 0
urlRegex
string

Regex pattern. Only URLs matching this pattern will be followed and scraped. An automatic prefix scope in the form ^ follows a redirect of the starting page.

Example:

"^https?://[^/]+/blog/"

Preserve hyperlinks in the Markdown output

includeImages
boolean
default:false

Include image references in the Markdown output

shortenBase64Images
boolean
default:true

Truncate base64-encoded image data in the Markdown output

useMainContentOnly
boolean
default:false

Extract only the main content, stripping headers, footers, sidebars, and navigation

followSubdomains
boolean
default:false

When true, follow links on subdomains of the starting URL's domain (e.g. docs.example.com when starting from example.com). www and apex are always treated as equivalent.

pdf
object

PDF handling. start/end limit parsing to an inclusive, 1-based page range.

includeFrames
boolean
default:false

When true, the contents of iframes are rendered to Markdown for each crawled page.

includeSelectors
string[]

Keep matching HTML subtrees before converting each page to Markdown.

Maximum array length: 50
Maximum string length: 2048
excludeSelectors
string[]

Remove matching elements after inclusions. Exclusions take precedence.

Maximum array length: 50
Maximum string length: 2048
maxAgeMs
integer
default:86400000

Maximum cache age in milliseconds. Defaults to 1 day; 0 fetches fresh.

Required range: 0 <= x <= 2592000000
waitForMs
integer
default:3500

Browser wait time in milliseconds after initial page load for each crawled page. Defaults to 3500 (3.5 seconds). Min: 0. Max: 30000 (30 seconds).

Required range: 0 <= x <= 30000
settleAnimations
boolean
default:false

Wait briefly for CSS animations and transitions to settle before reading each page.

stopAfterMs
integer
default:80000

Soft crawl deadline in milliseconds. Returns pages collected before the next deadline check.

Required range: 10000 <= x <= 110000
country
enum<string>

Fetch from this country (ISO 3166-1 alpha-2).

Available options:
ad,
ae,
af,
ag,
ai,
al,
am,
ao,
ar,
at,
au,
aw,
az,
ba,
bb,
bd,
be,
bf,
bg,
bh,
bi,
bj,
bm,
bn,
bo,
bq,
br,
bs,
bw,
by,
bz,
ca,
cd,
cf,
cg,
ch,
ci,
cl,
cm,
cn,
co,
cr,
cv,
cw,
cy,
cz,
de,
dj,
dk,
dm,
do,
dz,
ec,
ee,
eg,
es,
et,
fi,
fj,
fr,
ga,
gb,
gd,
ge,
gf,
gg,
gh,
gm,
gn,
gp,
gq,
gr,
gt,
gu,
gw,
gy,
hk,
hn,
hr,
ht,
hu,
id,
ie,
il,
im,
in,
iq,
ir,
is,
it,
je,
jm,
jo,
jp,
ke,
kg,
kh,
kn,
kr,
kw,
ky,
kz,
la,
lb,
lc,
lk,
lr,
ls,
lt,
lu,
lv,
ly,
ma,
mc,
md,
me,
mf,
mg,
mk,
ml,
mm,
mn,
mo,
mq,
mr,
mt,
mu,
mv,
mw,
mx,
my,
mz,
na,
nc,
ne,
ng,
ni,
nl,
no,
np,
nz,
om,
pa,
pe,
pf,
pg,
ph,
pk,
pl,
pr,
ps,
pt,
py,
qa,
re,
ro,
rs,
ru,
rw,
sa,
sc,
sd,
se,
sg,
si,
sk,
sl,
sm,
sn,
so,
sr,
ss,
st,
sv,
sx,
sy,
sz,
tc,
td,
tg,
th,
tj,
tl,
tm,
tn,
tr,
tt,
tw,
tz,
ua,
ug,
us,
uy,
uz,
vc,
ve,
vg,
vi,
vn,
ye,
yt,
za,
zm,
zw
Example:

"de"

timeoutOpts
object

Request deadline and what to return when it passes.

zdr
enum<string>
default:disabled

enabled turns on zero data retention. Returns 403 ZDR_NOT_ENABLED unless your organization has ZDR.

Available options:
enabled,
disabled
tags
string[]

Labels for filtering usage in the dashboard.

Maximum array length: 20
Required string length: 1 - 50
Example:

Response

Successful response

results
object[]
required
metadata
object
required
request_id
string<uuid>
required

Unique ID of this request, also in X-Request-Id. Include it when contacting support.

Example:

"3f1c2a6e-8b4d-4c1e-9f0a-2d7b5e6c8a91"

cache_metadata
object
required

Whether this response came from cache.

partial
boolean

True when timeoutOpts.behavior=return-partial returned the usable results collected before the deadline. Partial collections are not cached as complete results.

key_metadata
object

Credits this request used and your remaining balance.