Wykres commitów

1372 Commity (79f42c3c41d6f2127ef18b26c2e32f6ab211e00f)

Autor SHA1 Wiadomość Data
Miguel Sozinho Ramalho 79f42c3c41
Merge pull request #318 from bellingcat/feat/antibot-reddit
Adds RedditDropin and other flow improvements
2025-06-10 18:39:34 +01:00
msramalho 8314833ae8
removes exclude_media_extensions option 2025-06-10 18:34:33 +01:00
msramalho 6279610a43
updates docs 2025-06-10 18:28:45 +01:00
msramalho fc89d96517
escape sequence 2025-06-10 18:04:33 +01:00
msramalho 54fda9cad4
antibot in docker uses a different user_data_dir 2025-06-10 18:04:27 +01:00
msramalho 71636233cb
adds migration information and VkDropin info. 2025-06-10 17:07:10 +01:00
msramalho fdbe96f2e4
vk and reddit should work without credentials but log the error 2025-06-10 16:44:14 +01:00
msramalho 22bd8727df
python dependencies bump 2025-06-10 16:43:55 +01:00
msramalho 499c272260
dependabot switch to monthly 2025-06-10 16:37:52 +01:00
Miguel Sozinho Ramalho f232bc45b8
Merge pull request #315 from bellingcat/dependabot/docker/webrecorder/browsertrix-crawler-1.6.2
Bump webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2
2025-06-10 16:34:30 +01:00
msramalho 4270e06728
npm update on scripts/settings 2025-06-10 16:33:47 +01:00
msramalho ca00aa302d
version bump breaking 2025-06-10 16:31:32 +01:00
msramalho 773fa82f06
introduces reddit dropin 2025-06-10 16:31:19 +01:00
msramalho ef0e909a72
extractor to auto detect best quality 2025-06-10 16:29:35 +01:00
msramalho 6bbc7fb47a
improves antibot flow and makes auth_wall detection optional 2025-06-10 16:29:07 +01:00
msramalho 809b8c7749
default dropin introduced 2025-06-10 16:14:42 +01:00
msramalho 6d82655cc4
manifest improvement for antibot 2025-06-10 16:14:34 +01:00
msramalho 6bd493a791
dropin with new ytdlp feature and helper method 2025-06-10 16:11:55 +01:00
msramalho 287e823f43
improves twitter URL cleaning and introduces another bestquality check 2025-06-10 16:09:38 +01:00
msramalho c815488daa
adds new URLs to ignore 2025-06-10 15:44:52 +01:00
dependabot[bot] f53e34d6bd
Bump webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2
Bumps webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2.

---
updated-dependencies:
- dependency-name: webrecorder/browsertrix-crawler
  dependency-version: 1.6.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
2025-06-09 20:55:07 +00:00
Miguel Sozinho Ramalho 4cfbc3008b
Merge pull request #313 from bellingcat/feat/antibot-auth
Introduces more flexibility to the Antibot Extractor
2025-06-08 14:42:35 +01:00
msramalho 6f02493ff1
adds clips extraction to VK, though generic_extractor should still be run for those 2025-06-08 14:36:55 +01:00
msramalho 1f2d637928
minor improvements 2025-06-08 14:16:21 +01:00
msramalho 18cc05a2fe
allows auth_for_site to receive do.main directly 2025-06-08 14:16:12 +01:00
msramalho c96fd71f35
minor cleanup 2025-06-07 20:06:53 +01:00
msramalho b3183510ea
installs ffmpeg in GH actions 2025-06-07 20:03:26 +01:00
msramalho d13a5ef003
adds tests in minor improvements 2025-06-07 19:58:18 +01:00
msramalho 48c1ab3c1f
doc improvements 2025-06-07 19:14:16 +01:00
msramalho b2ee42ee95
adds the first antibot dropin: VKontakte 2025-06-07 19:10:01 +01:00
msramalho 07ff5baf07
adds Dropin flexible integration for antibot 2025-06-07 19:09:37 +01:00
msramalho d202d79e0f
lint 2025-06-07 19:06:14 +01:00
msramalho e2e6490b49
minimal changes 2025-06-07 18:15:21 +01:00
msramalho 952487da30
adds missing bin dependency 2025-06-07 18:14:42 +01:00
msramalho c7a84bc97a
generalizes ydl info to filename method for reusing 2025-06-07 18:14:08 +01:00
msramalho c0be41950d
Merge branch 'dev' of https://github.com/bellingcat/auto-archiver into dev 2025-06-04 17:06:42 +01:00
Miguel Sozinho Ramalho ae547ef83f
Merge pull request #308 from bellingcat/dependabot/npm_and_yarn/scripts/settings/actions-a541a3dacb
Bump the actions group in /scripts/settings with 4 updates
2025-06-04 15:06:59 +01:00
msramalho 8a897cf601
minimal changes: standard naming 2025-06-04 15:06:08 +01:00
Miguel Sozinho Ramalho 14c8af5cc8
Merge pull request #310 from djhmateer/waczscreenshot bug fix
counter_screenshots to counter_warc_files in wacz_extractor so don't …
2025-06-04 15:01:12 +01:00
Miguel Sozinho Ramalho 8e2e18ef75
Merge pull request #311 from bellingcat/feat/seleniumbase
Replaces ScreenshotEnricher with AntibotExtractorEnricher, removes VkExtractor
2025-06-04 14:53:31 +01:00
msramalho 5491f3e9e7
fixing s3 storage tests 2025-06-04 14:41:00 +01:00
msramalho 264ba82ea0
finish removing screenshot_enricher references 2025-06-04 14:31:07 +01:00
msramalho 05231445d9
removes unnecessary ignored files 2025-06-04 14:19:25 +01:00
msramalho 2c6be4447f
linting 2025-06-04 14:17:38 +01:00
msramalho 5f68c151a0
removes webdriver utils used by screenshot enricher 2025-06-04 14:17:19 +01:00
msramalho 6d2aec032f
Merge remote-tracking branch 'origin/main' into dev 2025-06-04 14:15:14 +01:00
msramalho bc8cf2fb29
minor TODO 2025-06-04 14:10:19 +01:00
msramalho f066111d49
removes geckodriver dependencies following screenshot enricher removal 2025-06-04 14:09:13 +01:00
msramalho e6f3826a3a
dropping screenshot enricher 2025-06-04 12:08:59 +01:00
msramalho e5a78a5d06
antibot can be used out of the box 2025-06-04 12:01:42 +01:00