mirror of https://github.com/webrecorder/browsertrix-crawler.git synced 2025-10-19 06:23:16 +00:00

Run a high-fidelity browser-based web archiving crawler in a single Docker container https://crawler.docs.browsertrix.com

crawler crawling wacz warc web-archiving web-crawler webrecorder

Find a file

Ilya Kreymer e5bab8e7c8 various edge-case loading optimizations: (#709 ) - rework 'should stream' logic: * ensure 206 responses (or any response) greater than 25M are streamed * response between 5M and 25M are read into memory if text/css/js as they may be rewritten * responses <5M are read into memory * responses with unknown size are streamed if a 2xx, otherwise read into memory, assuming error code responses may lack status codes but otherwise are small - likely fix for issues in #706 - if too many range requests for same URL are being made, try skipping/failing right away to reduce load - assume main browser context is used not just for service workers, always enable - check false positive 'net-aborted' error that may actually be ok for media, as well as documents - improve logging - interrupt any pending requests (that may be loading via browser context) after page timeout, log dropped requests --------- Co-authored-by: Tessa Walsh <tessa@bitarchivist.net>		2024-10-31 14:06:17 -07:00
.github/workflows	tests: disable blockrules youtube tests in CI (#698 )	2024-10-04 17:37:13 -07:00
.husky	Add MKDocs documentation site for Browsertrix Crawler 1.0.0 (#494 )	2024-03-16 14:59:32 -07:00
config/policies	cleanup: remove old config files from pywb (#682 )	2024-09-05 20:23:34 -07:00
docs	Add documentation for crawl collections (#695 )	2024-10-05 11:51:32 -07:00
html	Adblock support (#534 )	2024-04-12 09:47:32 -07:00
src	various edge-case loading optimizations: (#709 )	2024-10-31 14:06:17 -07:00
tests	tests: use old.webrecorder.net for testing (#710 )	2024-10-31 13:24:58 -04:00
.dockerignore	Add ad blocking via request interception (#173 )	2022-11-15 18:30:27 -08:00
.eslintignore	follow-up to #428 : update ignore files (#431 )	2023-11-09 17:13:53 -08:00
.eslintrc.cjs	eslint: add strict await checking: (#684 )	2024-09-06 16:24:18 -07:00
.gitignore	Gracefully handle non-absolute path for create-login-profile --filename (#521 )	2024-03-29 13:46:54 -07:00
.pre-commit-config.yaml	Add Prettier to the repo, and format all the files! (#428 )	2023-11-09 16:11:11 -08:00
.prettierignore	follow-up to #428 : update ignore files (#431 )	2023-11-09 17:13:53 -08:00
CHANGES.md	Add Prettier to the repo, and format all the files! (#428 )	2023-11-09 16:11:11 -08:00
docker-compose.yml	Add Prettier to the repo, and format all the files! (#428 )	2023-11-09 16:11:11 -08:00
docker-entrypoint.sh	Improved support for running as non-root (#503 )	2024-03-21 08:16:59 -07:00
Dockerfile	bump browser to 1.69.162 (#681 )	2024-09-05 20:21:43 -07:00
LICENSE	initial commit after split from zimit	2020-10-31 13:16:37 -07:00
NOTICE	initial commit after split from zimit	2020-10-31 13:16:37 -07:00
package.json	various edge-case loading optimizations: (#709 )	2024-10-31 14:06:17 -07:00
README.md	Add MKDocs documentation site for Browsertrix Crawler 1.0.0 (#494 )	2024-03-16 14:59:32 -07:00
requirements.txt	Separate writing pages to pages.jsonl + extraPages.jsonl to use with new py-wacz (#535 )	2024-04-11 13:55:52 -07:00
test-setup.js	Fix disk utilization computation errors (#338 )	2023-07-05 21:58:28 -07:00
tsconfig.eslint.json	eslint: add strict await checking: (#684 )	2024-09-06 16:24:18 -07:00
tsconfig.json	dep: update to wabac.js 2.20 (#704 )	2024-10-16 21:02:04 -07:00
yarn.lock	various edge-case loading optimizations: (#709 )	2024-10-31 14:06:17 -07:00

README.md

Browsertrix Crawler 1.x

Browsertrix Crawler is a standalone browser-based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. Browsertrix Crawler uses Puppeteer to control one or more Brave Browser browser windows in parallel. Data is captured through the Chrome Devtools Protocol (CDP) in the browser.

For information on how to use and develop Browsertrix Crawler, see the hosted Browsertrix Crawler documentation.

For information on how to build the docs locally, see the docs page.

Support

Initial support for 0.x version of Browsertrix Crawler, was provided by Kiwix. The initial functionality for Browsertrix Crawler was developed to support the zimit project in a collaboration between Webrecorder and Kiwix, and this project has been split off from Zimit into a core component of Webrecorder.

Additional support for Browsertrix Crawler, including for the development of the 0.4.x version has been provided by Portico.

License

AGPLv3 or later, see LICENSE for more details.