Skip to main content
Version: Next

CrawlingRequest <UserData>

A Request as seen by a crawler: the same stored record, with the crawler's own per-request state (userData.__crawlee) exposed as properties. This is what request is in every request handler.

Requests fetched from a request manager are plain Requests; BasicCrawler rebuilds them as CrawlingRequests before handing them to the pipeline.

Hierarchy

Index

Constructors

constructor

  • Request parameters including the URL, HTTP method and headers, and others.

    Processing state (retryCount, errorMessages, handledAt, loadedUrl, id) is not an option; requests coming back from a storage are rebuilt with Request.fromSchema.


    Parameters

    Returns CrawlingRequest<UserData>

Properties

inheritederrorMessages

errorMessages: string[]

An array of error messages from request processing.

optionalinheritedhandledAt

handledAt?: string

ISO datetime string that indicates the time when the request has been processed. Is null if the request has not been crawled yet.

optionalinheritedheaders

headers?: Record<string, string>

Object with HTTP headers. Key is header name, value is the value.

optionalinheritedid

id?: string

Storage-assigned request ID. Only present on requests that went through a RequestQueue.

optionalinheritedloadedUrl

loadedUrl?: string

An actually loaded URL after redirects, if present. HTTP redirects are guaranteed to be included.

When using PuppeteerCrawler or PlaywrightCrawler, meta tag and JavaScript redirects may, or may not be included, depending on their nature. This generally means that redirects, which happen immediately will most likely be included, but delayed redirects will not.

inheritedmethod

HTTP method, e.g. GET or POST.

inheritednoRetry

noRetry: boolean

The true value indicates that the request will not be automatically retried on error.

optionalinheritedpayload

payload?: string

HTTP request payload, e.g. for POST requests.

inheritedretryCount

retryCount: number

Indicates the number of times the crawling of the request has been retried on error.

inheriteduniqueKey

uniqueKey: string

A unique key identifying the request. Two requests with the same uniqueKey are considered as pointing to the same URL.

inheritedurl

url: string

URL of the web page to crawl.

inheriteduserData

userData: UserData = ...

Custom user data assigned to the request.

All data stored in userData must be JSON-serializable. Storing non-serializable values (e.g. functions, symbols) may result in unexpected results.

Accessors

crawlDepth

  • get crawlDepth(): number
  • set crawlDepth(value): void
  • Depth of the request in the current crawl tree. Note that this is dependent on the crawler setup and might produce unexpected results when used with multiple crawlers.


    Returns number

  • Parameters

    • value: number

    Returns void

inheritedlabel

  • get label(): undefined | string
  • set label(value): void
  • shortcut for getting request.userData.label


    Returns undefined | string

  • shortcut for setting request.userData.label


    Parameters

    • value: undefined | string

    Returns void

maxRetries

  • get maxRetries(): undefined | number
  • set maxRetries(value): void
  • Maximum number of retries for this request. Allows to override the global maxRequestRetries option of BasicCrawler.


    Returns undefined | number

  • Parameters

    • value: undefined | number

    Returns void

sessionId

  • get sessionId(): undefined | string
  • set sessionId(value): void
  • ID of a session to use for this request. When set, the crawler will fetch this session from the session pool instead of creating a new one.


    Returns undefined | string

  • Parameters

    • value: undefined | string

    Returns void

skipNavigation

  • get skipNavigation(): boolean
  • set skipNavigation(value): void
  • Tells the crawler processing this request to skip the navigation and process the request directly.

    When this is set to true, the crawling context will not contain the results of the navigation (e.g. response, body, contentType, $ or request.loadedUrl). Accessing these properties will throw a NavigationSkippedError at runtime.


    Returns boolean

  • Parameters

    • value: boolean

    Returns void

state

Methods

publicinheritedintoFetchAPIRequest

  • intoFetchAPIRequest(): Request
  • Converts the Crawlee Request object to a fetch API Request object.


    Returns Request

    The native fetch API Request object.

pushErrorMessage

  • pushErrorMessage(errorOrMessage, options): void
  • Stores information about an error that occurred during processing of this request.

    You should always use Error instances when throwing errors in JavaScript.

    Nevertheless, to improve the debugging experience when using third party libraries that may not always throw an Error instance, the function performs a type inspection of the passed argument and attempts to extract as much information as possible, since just throwing a bad type error makes any debugging rather difficult.


    Parameters

    • errorOrMessage: unknown

      Error object or error message to be stored in the request.

    • optionaloptions: PushErrorMessageOptions = {}

    Returns void

staticfromSchema

  • Rebuilds a request from its stored form, including the processing state a RequestOptions object cannot carry. Subclasses get instances of themselves.


    Parameters

    Returns CrawlingRequest<UserData>