Skip to main content
Version: Next

Request <UserData>

Represents a URL to be crawled, optionally including HTTP method, headers, payload and other metadata, together with the processing state a request queue keeps about it (retryCount, errorMessages, handledAt, ...). Inside a crawler, requests are CrawlingRequests, which add the crawler's own per-request settings.

Each Request instance has the uniqueKey property, which can be either specified manually in the constructor or generated automatically from the URL. Two requests with the same uniqueKey are considered as pointing to the same web resource. This behavior applies to all Crawlee classes, such as RequestList, RequestQueue, PuppeteerCrawler or PlaywrightCrawler.

To access and examine the actual request sent over http, with all autofilled headers you can access response.request object from the request handler

Example use:

const request = new Request({
url: 'http://www.example.com',
headers: { Accept: 'application/json' },
userData: { foo: 'bar' },
});

await requestQueue.addRequest(request);

Hierarchy

Index

Constructors

constructor

  • Request parameters including the URL, HTTP method and headers, and others.

    Processing state (retryCount, errorMessages, handledAt, loadedUrl, id) is not an option; requests coming back from a storage are rebuilt with Request.fromSchema.


    Parameters

    Returns CrawleeRequest<UserData>

Properties

errorMessages

errorMessages: string[]

An array of error messages from request processing.

optionalhandledAt

handledAt?: string

ISO datetime string that indicates the time when the request has been processed. Is null if the request has not been crawled yet.

optionalheaders

headers?: Record<string, string>

Object with HTTP headers. Key is header name, value is the value.

optionalid

id?: string

Storage-assigned request ID. Only present on requests that went through a RequestQueue.

optionalloadedUrl

loadedUrl?: string

An actually loaded URL after redirects, if present. HTTP redirects are guaranteed to be included.

When using PuppeteerCrawler or PlaywrightCrawler, meta tag and JavaScript redirects may, or may not be included, depending on their nature. This generally means that redirects, which happen immediately will most likely be included, but delayed redirects will not.

method

HTTP method, e.g. GET or POST.

noRetry

noRetry: boolean

The true value indicates that the request will not be automatically retried on error.

optionalpayload

payload?: string

HTTP request payload, e.g. for POST requests.

retryCount

retryCount: number

Indicates the number of times the crawling of the request has been retried on error.

uniqueKey

uniqueKey: string

A unique key identifying the request. Two requests with the same uniqueKey are considered as pointing to the same URL.

url

url: string

URL of the web page to crawl.

userData

userData: UserData = ...

Custom user data assigned to the request.

All data stored in userData must be JSON-serializable. Storing non-serializable values (e.g. functions, symbols) may result in unexpected results.

Accessors

label

  • get label(): undefined | string
  • set label(value): void
  • shortcut for getting request.userData.label


    Returns undefined | string

  • shortcut for setting request.userData.label


    Parameters

    • value: undefined | string

    Returns void

Methods

publicintoFetchAPIRequest

  • intoFetchAPIRequest(): Request
  • Converts the Crawlee Request object to a fetch API Request object.


    Returns Request

    The native fetch API Request object.

staticfromSchema

  • Rebuilds a request from its stored form, including the processing state a RequestOptions object cannot carry. Subclasses get instances of themselves.


    Parameters

    Returns CrawleeRequest<UserData>