Miguel Sozinho Ramalho
|
79f42c3c41
|
Merge pull request #318 from bellingcat/feat/antibot-reddit
Adds RedditDropin and other flow improvements
|
2025-06-10 18:39:34 +01:00 |
msramalho
|
8314833ae8
|
removes exclude_media_extensions option
|
2025-06-10 18:34:33 +01:00 |
msramalho
|
6279610a43
|
updates docs
|
2025-06-10 18:28:45 +01:00 |
msramalho
|
fc89d96517
|
escape sequence
|
2025-06-10 18:04:33 +01:00 |
msramalho
|
54fda9cad4
|
antibot in docker uses a different user_data_dir
|
2025-06-10 18:04:27 +01:00 |
msramalho
|
71636233cb
|
adds migration information and VkDropin info.
|
2025-06-10 17:07:10 +01:00 |
msramalho
|
fdbe96f2e4
|
vk and reddit should work without credentials but log the error
|
2025-06-10 16:44:14 +01:00 |
msramalho
|
22bd8727df
|
python dependencies bump
|
2025-06-10 16:43:55 +01:00 |
msramalho
|
499c272260
|
dependabot switch to monthly
|
2025-06-10 16:37:52 +01:00 |
Miguel Sozinho Ramalho
|
f232bc45b8
|
Merge pull request #315 from bellingcat/dependabot/docker/webrecorder/browsertrix-crawler-1.6.2
Bump webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2
|
2025-06-10 16:34:30 +01:00 |
msramalho
|
4270e06728
|
npm update on scripts/settings
|
2025-06-10 16:33:47 +01:00 |
msramalho
|
ca00aa302d
|
version bump breaking
|
2025-06-10 16:31:32 +01:00 |
msramalho
|
773fa82f06
|
introduces reddit dropin
|
2025-06-10 16:31:19 +01:00 |
msramalho
|
ef0e909a72
|
extractor to auto detect best quality
|
2025-06-10 16:29:35 +01:00 |
msramalho
|
6bbc7fb47a
|
improves antibot flow and makes auth_wall detection optional
|
2025-06-10 16:29:07 +01:00 |
msramalho
|
809b8c7749
|
default dropin introduced
|
2025-06-10 16:14:42 +01:00 |
msramalho
|
6d82655cc4
|
manifest improvement for antibot
|
2025-06-10 16:14:34 +01:00 |
msramalho
|
6bd493a791
|
dropin with new ytdlp feature and helper method
|
2025-06-10 16:11:55 +01:00 |
msramalho
|
287e823f43
|
improves twitter URL cleaning and introduces another bestquality check
|
2025-06-10 16:09:38 +01:00 |
msramalho
|
c815488daa
|
adds new URLs to ignore
|
2025-06-10 15:44:52 +01:00 |
dependabot[bot]
|
f53e34d6bd
|
Bump webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2
Bumps webrecorder/browsertrix-crawler from 1.6.1 to 1.6.2.
---
updated-dependencies:
- dependency-name: webrecorder/browsertrix-crawler
dependency-version: 1.6.2
dependency-type: direct:production
update-type: version-update:semver-patch
...
Signed-off-by: dependabot[bot] <support@github.com>
|
2025-06-09 20:55:07 +00:00 |
Miguel Sozinho Ramalho
|
4cfbc3008b
|
Merge pull request #313 from bellingcat/feat/antibot-auth
Introduces more flexibility to the Antibot Extractor
|
2025-06-08 14:42:35 +01:00 |
msramalho
|
6f02493ff1
|
adds clips extraction to VK, though generic_extractor should still be run for those
|
2025-06-08 14:36:55 +01:00 |
msramalho
|
1f2d637928
|
minor improvements
|
2025-06-08 14:16:21 +01:00 |
msramalho
|
18cc05a2fe
|
allows auth_for_site to receive do.main directly
|
2025-06-08 14:16:12 +01:00 |
msramalho
|
c96fd71f35
|
minor cleanup
|
2025-06-07 20:06:53 +01:00 |
msramalho
|
b3183510ea
|
installs ffmpeg in GH actions
|
2025-06-07 20:03:26 +01:00 |
msramalho
|
d13a5ef003
|
adds tests in minor improvements
|
2025-06-07 19:58:18 +01:00 |
msramalho
|
48c1ab3c1f
|
doc improvements
|
2025-06-07 19:14:16 +01:00 |
msramalho
|
b2ee42ee95
|
adds the first antibot dropin: VKontakte
|
2025-06-07 19:10:01 +01:00 |
msramalho
|
07ff5baf07
|
adds Dropin flexible integration for antibot
|
2025-06-07 19:09:37 +01:00 |
msramalho
|
d202d79e0f
|
lint
|
2025-06-07 19:06:14 +01:00 |
msramalho
|
e2e6490b49
|
minimal changes
|
2025-06-07 18:15:21 +01:00 |
msramalho
|
952487da30
|
adds missing bin dependency
|
2025-06-07 18:14:42 +01:00 |
msramalho
|
c7a84bc97a
|
generalizes ydl info to filename method for reusing
|
2025-06-07 18:14:08 +01:00 |
msramalho
|
c0be41950d
|
Merge branch 'dev' of https://github.com/bellingcat/auto-archiver into dev
|
2025-06-04 17:06:42 +01:00 |
Miguel Sozinho Ramalho
|
ae547ef83f
|
Merge pull request #308 from bellingcat/dependabot/npm_and_yarn/scripts/settings/actions-a541a3dacb
Bump the actions group in /scripts/settings with 4 updates
|
2025-06-04 15:06:59 +01:00 |
msramalho
|
8a897cf601
|
minimal changes: standard naming
|
2025-06-04 15:06:08 +01:00 |
Miguel Sozinho Ramalho
|
14c8af5cc8
|
Merge pull request #310 from djhmateer/waczscreenshot bug fix
counter_screenshots to counter_warc_files in wacz_extractor so don't …
|
2025-06-04 15:01:12 +01:00 |
Miguel Sozinho Ramalho
|
8e2e18ef75
|
Merge pull request #311 from bellingcat/feat/seleniumbase
Replaces ScreenshotEnricher with AntibotExtractorEnricher, removes VkExtractor
|
2025-06-04 14:53:31 +01:00 |
msramalho
|
5491f3e9e7
|
fixing s3 storage tests
|
2025-06-04 14:41:00 +01:00 |
msramalho
|
264ba82ea0
|
finish removing screenshot_enricher references
|
2025-06-04 14:31:07 +01:00 |
msramalho
|
05231445d9
|
removes unnecessary ignored files
|
2025-06-04 14:19:25 +01:00 |
msramalho
|
2c6be4447f
|
linting
|
2025-06-04 14:17:38 +01:00 |
msramalho
|
5f68c151a0
|
removes webdriver utils used by screenshot enricher
|
2025-06-04 14:17:19 +01:00 |
msramalho
|
6d2aec032f
|
Merge remote-tracking branch 'origin/main' into dev
|
2025-06-04 14:15:14 +01:00 |
msramalho
|
bc8cf2fb29
|
minor TODO
|
2025-06-04 14:10:19 +01:00 |
msramalho
|
f066111d49
|
removes geckodriver dependencies following screenshot enricher removal
|
2025-06-04 14:09:13 +01:00 |
msramalho
|
e6f3826a3a
|
dropping screenshot enricher
|
2025-06-04 12:08:59 +01:00 |
msramalho
|
e5a78a5d06
|
antibot can be used out of the box
|
2025-06-04 12:01:42 +01:00 |