Proxy Management
IP address blocking is one of the oldest and most effective ways of preventing access to a website. It is therefore paramount for a good web scraping library to provide easy to use but powerful tools which can work around IP blocking. The most powerful weapon in our anti IP blocking arsenal is a proxy server.
With Crawlee we can use our own proxy servers or proxy servers acquired from third-party providers.
Check out the avoid blocking guide for more information about blocking.
Quick start
If we already have proxy URLs of our own, we can start using them immediately in only a few lines of code.
import { ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: [
'http://proxy-1.com',
'http://proxy-2.com',
]
});
const proxyUrl = await proxyConfiguration.newUrl();
Examples of how to use our proxy URLs with crawlers are shown below in Crawler integration section.
Proxy Configuration
All our proxy needs are managed by the ProxyConfiguration class. We create an instance using the ProxyConfiguration constructor function based on the provided options. See the ProxyConfigurationOptions for all the possible constructor options.
Static proxy list
You can provide a static list of proxy URLs to the proxyUrls option. The ProxyConfiguration will then rotate through the provided proxies.
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: [
'http://proxy-1.com',
'http://proxy-2.com',
null // null means no proxy is used
]
});
This is the simplest way to use a list of proxies. Crawlee will rotate through the list of proxies in a round-robin fashion.
Custom proxy function
The ProxyConfiguration class allows you to provide a custom function to pick a proxy URL. This is useful when you want to implement your own logic for selecting a proxy.
const proxyConfiguration = new ProxyConfiguration({
newUrlFunction: async () => {
const { proxyUrl } = await myProxyProvider.getProxy();
return proxyUrl; // or return null to skip the proxy
}
});
The newUrlFunction takes no arguments and returns a string with the proxy URL (or null to skip the proxy). Crawlers call it once for every new session, so it cannot route individual requests. To send specific requests through specific proxies, create named sessions with their own proxies and pin requests to them, as shown in the session management guide.
Crawler integration
ProxyConfiguration integrates seamlessly into HttpCrawler, CheerioCrawler, PlaywrightCrawler and PuppeteerCrawler.
- HttpCrawler
- CheerioCrawler
- PlaywrightCrawler
- PuppeteerCrawler
import { HttpCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: ['http://proxy-1.com', 'http://proxy-2.com'],
});
const crawler = new HttpCrawler({
proxyConfiguration,
// ...
});
import { CheerioCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: ['http://proxy-1.com', 'http://proxy-2.com'],
});
const crawler = new CheerioCrawler({
proxyConfiguration,
// ...
});
import { PlaywrightCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: ['http://proxy-1.com', 'http://proxy-2.com'],
});
const crawler = new PlaywrightCrawler({
proxyConfiguration,
// ...
});
import { PuppeteerCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: ['http://proxy-1.com', 'http://proxy-2.com'],
});
const crawler = new PuppeteerCrawler({
proxyConfiguration,
// ...
});
Our crawlers will now use the selected proxies for all connections.
IP Rotation and session management
Each call to proxyConfiguration.newUrl() generates a new proxy URL. Crawler instances pair these URLs with Session instances and rotate those together with browser fingerprints, impersonated headers, and more. This is extremely useful in scraping, because we want to create the impression of a real user. See the session management guide and SessionPool class for more information on how keeping a real session helps us avoid blocking.
- HttpCrawler
- CheerioCrawler
- PlaywrightCrawler
- PuppeteerCrawler
- Standalone
import { HttpCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new HttpCrawler({
saveResponseCookies: true,
proxyConfiguration,
// ...
});
import { CheerioCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new CheerioCrawler({
saveResponseCookies: true,
proxyConfiguration,
// ...
});
import { PlaywrightCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new PlaywrightCrawler({
saveResponseCookies: true,
proxyConfiguration,
// ...
});
import { PuppeteerCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new PuppeteerCrawler({
saveResponseCookies: true,
proxyConfiguration,
// ...
});
import { ProxyConfiguration, SessionPool } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const proxyUrl = await proxyConfiguration.newUrl();
Inspecting current proxy in Crawlers
HttpCrawler, CheerioCrawler, PlaywrightCrawler and PuppeteerCrawler grant access to information about the currently used proxy
in their requestHandler using a proxyInfo object.
With the proxyInfo object, we can easily access the proxy URL.
- HttpCrawler
- CheerioCrawler
- PlaywrightCrawler
- PuppeteerCrawler
import { HttpCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new HttpCrawler({
proxyConfiguration,
async requestHandler({ proxyInfo }) {
console.log(proxyInfo);
},
// ...
});
import { CheerioCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new CheerioCrawler({
proxyConfiguration,
async requestHandler({ proxyInfo }) {
console.log(proxyInfo);
},
// ...
});
import { PlaywrightCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new PlaywrightCrawler({
proxyConfiguration,
async requestHandler({ proxyInfo }) {
console.log(proxyInfo);
},
// ...
});
import { PuppeteerCrawler, ProxyConfiguration } from 'crawlee';
const proxyConfiguration = new ProxyConfiguration({/* opts */});
const crawler = new PuppeteerCrawler({
proxyConfiguration,
async requestHandler({ proxyInfo }) {
console.log(proxyInfo);
},
// ...
});