Stowage/browsertrix-crawler - Remotebranch.eu

Stowage/browsertrix-crawler

mirror of https://github.com/webrecorder/browsertrix-crawler.git synced 2025-10-19 14:33:17 +00:00

Author	SHA1	Message	Date
Ilya Kreymer	c0508d44a7	add locking mechanism for deleteOnExit and runnning post-crawl steps only run postcrawl steps on last instance, when all locks released	2022-03-16 04:16:33 +00:00
Ilya Kreymer	761ce7067b	behaviors update (#105 ) * update to browsertrix-behaviors 0.2.5 to support improved autoscroll - add evaluateWithCLI() to support evaluate() with 'getEventListeners()' and other devtools command-line api functions, to allow autoscroll behavior to check if it should exit out early - inject behaviors into interactive loader to allow testing - fix signal handler if state not inited yet - dependencies: update puppeteer-cluster to latest, update pywb to 2.6.5	2022-02-20 22:22:19 -08:00
Ilya Kreymer	39ddecd35e	State Save + Restore State from Config + Redis State + Scope Fix 0.5.0 (#78 ) * save state work: - support interrupting and saving crawl - support loading crawl state (frontier queue, pending, done) from YAML - support scope check when loading to apply new scoping rules when restarting crawl - failed urls added to done as failed, can be retried if crawl is stopped and restarted - save state to crawls/crawl-<ts>-<id>.yaml when interrupted - --saveState option controls when crawl state is saved, default to partial/when interrupted, also always, never. - support in-memory or redis based crawl state, using fork of puppeteer-cluster - --redisStore used to enable redis-based state * signals/crawl interruption: - crawl state set to drain/not provide any more urls to crawl - graceful stop of crawl in response to sigint/sigterm - initial sigint/sigterm waits for graceful end of current pages, second terminates immediately - initial sigabrt followed by sigterm terminates immediately - puppeteer disable handleSIGTERM, handleSIGHUP, handleSIGINT * redis state support: - use lua scripts for atomic move from queue -> pending, and pending -> done - pending key expiry set to page timeout - add numPending() and numSeen() to support better puppeteer-cluster semantics for early termination - drainMax returns the numPending() + numSeen() to work with cluster stats * arg improvements: - add --crawlId param, also settable via CRAWL_ID env var, defaulting to os.hostname() (used for redis key and crawl state file) - support setting cmdline args via env var CRAWL_ARGS - use 'choices' in args when possible * build update: - switch base browser image to new webrecorder/browsertrix-browser-base, simple image with .deb files only for amd64 and arm64 builds - use setuptools<58.0 * misc crawl/scoping rule fixes: - scoping rules fix when external is used with scopeType state: - limit: ensure no urls, including initial seeds, are added past the limit - signals: fix immediate shutdown on second signal - tests: add scope test for default scope + excludes * py-wacz update - add 'seed': true to pages that are seeds for optimized wacz creation, keeping non-seeds separate (supported via wacz 0.3.2) - pywb: use latest pywb branch for improved twitter video capture * update to latest browsertrix-behaviors * fix setuptools dependency #88 * update README for 0.5.0 beta	2021-09-28 09:41:16 -07:00
Ilya Kreymer	8f740d4e24	support custom crawl directory with --cwd flag, default to /crawls update README	2020-11-02 15:28:19 +00:00
Ilya Kreymer	a875aa90d3	Dockerfile: switch to cmd 'crawl', instead of entrypoint to support running 'pywb' also update README with docker-compose and docker run examples, update commandline example default output to './crawls' subdirectory	2020-11-01 21:35:00 -08:00
Ilya Kreymer	91b8994a08	refactor crawler and default driver: - add extensible defaultDriver, wrap crawling functionality in Crawler class - support headless/non-headless, custom driver - support custom collection name for pywb, generate-cdx option - autoplay: add slightly delay for splash loading	2020-11-01 19:53:47 -08:00