Upgrading to v4: removed symbols
The full list of removed exports and members, for ctrl-F purposes. Where a replacement exists, it is noted inline. This page is part of the Upgrading to v4 guide.
Removed symbols
BasicCrawler._cleanupContext(protected) - this is now handled by theContextPipelineBasicCrawler.isRequestBlocked(protected)BasicCrawler.events(protected) - this should be accessed viaBasicCrawler.serviceLocatorBrowserRequestHandlerandBrowserErrorHandlertypes in@crawlee/browserBrowserCrawler.userProvidedRequestHandler(protected)BrowserCrawler.requestHandlerTimeoutInnerMillis(protected)BrowserCrawler._enhanceCrawlingContextWithPageInfo(protected)BrowserCrawler._handleNavigation(protected)HttpCrawler.userRequestHandlerTimeoutMillis(protected)HttpCrawler._handleNavigation(protected)HttpCrawler._applyCookies(protected) - cookie merging is now handled byBaseHttpClientHttpCrawler._parseHTML(protected)HttpCrawler.useand theCrawlerExtensionclass (experimental) - theContextPipelineshould be used for extending the crawlerBasicCrawler._tagUserHandlerError(protected) - internal error-tagging helper, no longer part of the crawler surfaceBasicCrawler.handledRequestsCountsetter (@deprecated) - the throw-on-assign guard is gone; the getter is now internal-only and the count is derived fromthis.statisticsPlaywrightPlugin._containerProxyServer(public) - was an unused, never-populated fieldSnapshotter._snapshotMemory,Snapshotter._memoryOverloadWarning,Snapshotter._snapshotEventLoop,Snapshotter._snapshotCpu,Snapshotter._snapshotClient,Snapshotter._pruneSnapshots(all@deprecatedprotected stubs) - snapshotting is now handled by the individual load signals, and theSnapshotteritself is internal toConcurrencySystem; there is no longer a public API for reading raw resource snapshotsFileDownloadOptions.streamHandler- streaming should now be handled directly in therequestHandlerinsteadplaywrightUtils.registerUtilsToContextandpuppeteerUtils.registerUtilsToContext- this is now added to the context viaContextPipelinecompositioncontext.blockResourcesandcontext.cacheResponses, and thepuppeteerUtils.blockResources/puppeteerUtils.cacheResponsesfunctions behind them — both had a severe performance cost in recent Puppeteer versions and were already deprecated. UsepuppeteerUtils.blockRequests(page, options), which blocks URL patterns over CDP without disabling the browser cache. If you were caching responses, rely on the in-browser cache instead.context.closeCookieModals,playwrightUtils.closeCookieModalsandpuppeteerUtils.closeCookieModals— removed along with the optionalidcac-playwrightpeer dependency (see Crawling context no longer includescloseCookieModalsand the cookie modals guide)Configuration.systemInfoV2/CRAWLEE_SYSTEM_INFO_V2environment variable — the v2 behavior is now the default (see Available resource detection)KeyValueStore.getInput()andConfiguration.inputKey/CRAWLEE_INPUT_KEY— reading the run input moved to the Apify SDK (seeKeyValueStore.getInput()andConfiguration.inputKeymoved to the Apify SDK)Configuration.defaultDatasetId/defaultKeyValueStoreId/defaultRequestQueueIdand theirCRAWLEE_DEFAULT_*_IDenvironment variables — the default storage is addressed by a reserved alias, not by a configurable ID. Open a storage by name if you need a specific one.checkAndSerializeandchunkBySizefunctions (from@crawlee/core) — value (de)serialization now lives in theKeyValueStorefrontend; useserializeValue/parseValue(seemaybeStringifyis removed)BASIC_CRAWLER_TIMEOUT_BUFFER_SECSconstant (from@crawlee/basic) — was an internal timeout buffer, no longer exportedHttpResponse,HttpResponseWithoutBody,StreamingHttpResponse,ResponseTypes,BaseHttpResponseData,SimpleHeaders,processHttpRequestOptions, andGotScrapingHttpClient(from@crawlee/core) — the HTTP client surface moved to@crawlee/http-client/@crawlee/got-scraping-client(see HTTP client packages andBaseHttpClientreshaped)StreamHandlerContextandFileDownloadOptionstypes (from@crawlee/http) — seeFileDownloadnow extendsBasicCrawlerPlainResponsetype (from@crawlee/http) — it wrapped thegot-scrapingresponse and is gone along with the rest of the old HTTP response surface (seeCrawlingContext.responseis now of typeResponse)checkStorageAccess,withCheckedStorageAccessand theRequestHandlerResulttype — superseded by the storage transaction mechanism; usewithDirectStorageAccess()andStorageTransactionView(see Storage writes in request handlers are transactional)CreateContextOptionstype (from@crawlee/basic) — a leftover of the pre-ContextPipelinecontext-creation design, unused by the library itself; context construction is now driven byContextPipelineResponseLikeinterface (from@crawlee/core) — a vestige of the pre-fetchHTTP implementation with no consumers;getCookiesFromResponse()has always taken a nativeResponseUrlPatternObject(from@crawlee/core) — the compiled form of a URL pattern, produced internally by theenqueueLinks()machinery. Keep usingUrlPatternInput/GlobInput/RegExpInput, which are unchanged, and let the return type of the pattern helpers be inferredPERSIST_STATE_KEY(from@crawlee/core) — to change where a session pool persists its state, passpersistStateKeytoSessionPoolMAX_POOL_SIZEconstant (from@crawlee/core) — was the internal default forSessionPoolOptions.maxPoolSize(1000); inline the literal if you were reading itWithRequiredtype (from@crawlee/core) — a bare TypeScript utility that was never crawlee vocabulary;LoadedRequestno longer goes through it, so declare your own if you were using itErrorSnapshotterand itsSnapshotResultreturn type (from@crawlee/core) — an implementation detail ofErrorTracker. Error snapshotting is opt-in throughnew Statistics({ saveErrorSnapshots: true })(ornew ErrorTracker({ saveErrorSnapshots: true }))ErrorTracker.errorSnapshotterandErrorTracker.captureSnapshot()— both private now. Snapshotting is driven fromaddAsync()on the first occurrence of each distinct error; the captured URLs surface asfirstErrorScreenshotUrl/firstErrorHtmlUrlon the corresponding node oferrorTracker.result, as beforeMinimumSpeedStreamandByteCounterStream(from@crawlee/http) — theseTransformfactories existed only to be piped insideFileDownloadOptions.streamHandler, which v4 removed. Compose your ownTransformaroundcontext.response.bodyin therequestHandlerinstead; see the file download with streams exampleHttpHook,FileDownloadHook,CheerioHook,JSDOMHookandLinkeDOMHooktypes — see Removed navigation hook type aliases- The
puppeteerClickElementsnamespace (from@crawlee/puppeteer) —clickElements,clickElementsAndInterceptNavigationRequestsandisTargetRelevantwere internal helpers. UsepuppeteerUtils.enqueueLinksByClickingElements(), orcontext.enqueueLinksByClickingElements()inside a request handler; theEnqueueLinksByClickingElementsOptionstype is still exported directly from@crawlee/puppeteer - The
puppeteerRequestInterceptionnamespace (from@crawlee/puppeteer) — it only duplicatedpuppeteerUtils.addInterceptRequestHandler/puppeteerUtils.removeInterceptRequestHandler, which are unchanged. TheInterceptHandlertype is still exported directly from@crawlee/puppeteer - The top-level type exports
BlockRequestsOptions,InjectFileOptions,InfiniteScrollOptions,SaveSnapshotOptions,CompiledScriptParamsandCompiledScriptFunction(from@crawlee/puppeteer) — the types themselves are unchanged and remain reachable through the namespace, e.g.import { puppeteerUtils } from 'crawlee'; let options: puppeteerUtils.SaveSnapshotOptions;. This matches@crawlee/playwright, which never exported them at the top level.PuppeteerDirectNavigationOptionsis unaffected
The protected BasicCrawler.crawlingContexts map is removed
The property was not used by the library itself and re-implementing the functionality in user code is fairly straightforward.