Upgrading to v4: `@crawlee/utils` and `@crawlee/types`
Applies when you import utility functions, enums or types directly from @crawlee/utils or @crawlee/types, rather than only using the crawlers. This page is part of the Upgrading to v4 guide.
Available resource detection
In v3, we introduced a new way to detect available resources for the crawler, available via systemInfoV2 flag. In v4, this is the default way to detect available resources. The old way is removed completely together with the systemInfoV2 flag.
As part of this change, the low-level resource- and environment-detection helpers exported from @crawlee/utils were removed: getMemoryInfo() (and the MemoryInfo interface), isContainerized(), isDocker(), isLambda(), and getCgroupsVersion(). These backed the old detection path and are no longer part of the public API. Resource detection is now handled internally by the crawler's autoscaling; if you called any of these directly, read the equivalent values from the OS (node:os) or the relevant cgroup files yourself.
@crawlee/types symbols are no longer re-exported
The general-purpose utility types owned by @crawlee/types are no longer re-exported from other packages, so Dictionary, Awaitable, Constructor, Cookie, QueueOperationInfo and AllowedHttpMethods are no longer available from @crawlee/core (nor, in turn, from @crawlee/basic and the crawlee meta-package). Add @crawlee/types to your dependencies and import them from there — most of the package's types (ISession, ProxyInfo, RequestSchema, …) already required this. The interfaces you implement against — StorageBackend, StorageIdentifier and IBrowserPool / NewPageOptions — stay reachable from @crawlee/core and @crawlee/browser-pool respectively.
Removed and relocated @crawlee/utils exports
Besides the resource-detection helpers above, several other @crawlee/utils exports were removed or moved:
- Removed URL helpers:
filterUrl(target, origin, strategy),matchesEnqueueStrategy(strategy, target, origin), and theUNSUPPORTED_SCHEME_MESSAGEconstant. URL filtering by enqueue strategy is now internal toenqueueLinks. The relatedfilterRequestsByPatterns(requests, patterns?, onSkippedUrl?)function (from@crawlee/core) was removed for the same reason — pattern-based request filtering now happens insideenqueueLinks. - Relocated enums/types:
EnqueueStrategyis now exported from@crawlee/core,SearchParamsfrom@crawlee/types. They are no longer re-exported from@crawlee/utils, soimport { EnqueueStrategy } from '@crawlee/utils'breaks — import them fromcrawlee(the meta-package) or from@crawlee/core/@crawlee/typesinstead. - Removed
RobotsFilealias:RobotsFilewas an alias for theRobotsTxtFileclass and is removed. Rename any usage toRobotsTxtFile; the class itself is unchanged apart from the signature change described below. - Split into public and
/internalentry points: the main@crawlee/utilsentry now exposes only the user-facing helpers (sleep,htmlToText,extractUrls,downloadListOfUrls, thesocialnamespace, the Open Graph parser, and the robots/sitemap utilitiesRobotsTxtFile,SitemapanddiscoverValidSitemaps). Helpers that primarily serve the crawler packages - e.g.URL_NO_COMMAS_REGEX,URL_WITH_COMMAS_REGEX,extractUrlsFromCheerio,tryAbsoluteURL,expandShadowRoots, and the blocked-detection and iterable helpers - moved to the@crawlee/utils/internalentry point, which carries no semver guarantees. They keep working, but imports need updating:import { URL_NO_COMMAS_REGEX } from '@crawlee/utils/internal'. - Removed
CheerioRootand the cheerio type re-exports:CheerioRootwas an alias for cheerio's ownCheerioAPIand is gone;parseWithCheerio()andhtmlToText()are typed withCheerioAPIdirectly. The crawler packages also no longer re-exportCheerio,CheerioAPIandElement, soimport type { CheerioAPI } from 'crawlee'(or from@crawlee/basic/@crawlee/puppeteer/ ...) breaks - import them fromcheerioanddomhandler, which are the packages that own them. parseSitemapandexpandShadowRootsare no longer on the main entry:parseSitemap()(together with theSitemapUrltype) moved to@crawlee/utils/internal; use the documentedSitemap.load()/Sitemap.fromXmlString()/Sitemap.tryCommonNames()statics, ordiscoverValidSitemaps(), which stay on@crawlee/utils.expandShadowRoots()moved there too — it is a DOM function that is serialized into a browser page, not a Node helper. Because thecrawleemeta-package re-exports@crawlee/utilswholesale,import { parseSitemap } from 'crawlee'(and the same forexpandShadowRoots) breaks as well.
RobotsTxtFile.find signature changed; sitemap options removed
The proxyUrl argument of RobotsTxtFile.find() moved from a positional parameter into the options bag, which also gained httpClient and logger:
Before:
const robots = await RobotsTxtFile.find(url, proxyUrl, { timeoutMillis: 5000 });
After:
const robots = await RobotsTxtFile.find(url, { proxyUrl, timeoutMillis: 5000 });
Relatedly, the networkTimeouts option was dropped from ParseSitemapOptions; use the single timeoutMillis option instead. RobotsTxtFile.getSitemaps(), parseSitemaps() and parseUrlsFromSitemaps() still accept an optional RobotsTxtFileSitemapsOptions bag, whose enqueueStrategy option (default 'same-hostname') keeps only sitemap URLs on the robots.txt host — pass 'all' to disable that filtering. Non-http(s) sitemap URLs are always dropped.
HTML-parsing helper functions are now asynchronous
The HTML-parsing helper functions htmlToText, parseHandlesFromHtml and parseOpenGraph are now asynchronous and return promises.