Wykres commitów

1343 Commity (b2ee42ee9559d061fb0531828632b8a214d8eec4)

Autor SHA1 Wiadomość Data
msramalho b2ee42ee95
adds the first antibot dropin: VKontakte 2025-06-07 19:10:01 +01:00
msramalho 07ff5baf07
adds Dropin flexible integration for antibot 2025-06-07 19:09:37 +01:00
msramalho d202d79e0f
lint 2025-06-07 19:06:14 +01:00
msramalho e2e6490b49
minimal changes 2025-06-07 18:15:21 +01:00
msramalho 952487da30
adds missing bin dependency 2025-06-07 18:14:42 +01:00
msramalho c7a84bc97a
generalizes ydl info to filename method for reusing 2025-06-07 18:14:08 +01:00
msramalho c0be41950d
Merge branch 'dev' of https://github.com/bellingcat/auto-archiver into dev 2025-06-04 17:06:42 +01:00
Miguel Sozinho Ramalho ae547ef83f
Merge pull request #308 from bellingcat/dependabot/npm_and_yarn/scripts/settings/actions-a541a3dacb
Bump the actions group in /scripts/settings with 4 updates
2025-06-04 15:06:59 +01:00
msramalho 8a897cf601
minimal changes: standard naming 2025-06-04 15:06:08 +01:00
Miguel Sozinho Ramalho 14c8af5cc8
Merge pull request #310 from djhmateer/waczscreenshot bug fix
counter_screenshots to counter_warc_files in wacz_extractor so don't …
2025-06-04 15:01:12 +01:00
Miguel Sozinho Ramalho 8e2e18ef75
Merge pull request #311 from bellingcat/feat/seleniumbase
Replaces ScreenshotEnricher with AntibotExtractorEnricher, removes VkExtractor
2025-06-04 14:53:31 +01:00
msramalho 5491f3e9e7
fixing s3 storage tests 2025-06-04 14:41:00 +01:00
msramalho 264ba82ea0
finish removing screenshot_enricher references 2025-06-04 14:31:07 +01:00
msramalho 05231445d9
removes unnecessary ignored files 2025-06-04 14:19:25 +01:00
msramalho 2c6be4447f
linting 2025-06-04 14:17:38 +01:00
msramalho 5f68c151a0
removes webdriver utils used by screenshot enricher 2025-06-04 14:17:19 +01:00
msramalho 6d2aec032f
Merge remote-tracking branch 'origin/main' into dev 2025-06-04 14:15:14 +01:00
msramalho bc8cf2fb29
minor TODO 2025-06-04 14:10:19 +01:00
msramalho f066111d49
removes geckodriver dependencies following screenshot enricher removal 2025-06-04 14:09:13 +01:00
msramalho e6f3826a3a
dropping screenshot enricher 2025-06-04 12:08:59 +01:00
msramalho e5a78a5d06
antibot can be used out of the box 2025-06-04 12:01:42 +01:00
msramalho 258fb4faaf
visual HTML preview improvements 2025-06-04 12:00:40 +01:00
msramalho 5ec00f7811
adds dependencies for seleniumbase 2025-06-04 12:00:22 +01:00
msramalho 22408e2a98
adds test for antibot 2025-06-04 11:59:59 +01:00
msramalho 378b1a6d22
expand S3 objects content type for better preview results in non-latin languages 2025-06-04 11:53:41 +01:00
msramalho d130c1b3fa
WIP attempt at ytdlp impersonation 2025-06-04 11:53:18 +01:00
msramalho cbd189c97d
general cleanup 2025-06-04 11:53:01 +01:00
msramalho d2e8f1a512
introduces antibot step with seleniumbase 2025-06-04 11:20:46 +01:00
msramalho 488802b632
poetry update 2025-06-04 11:08:44 +01:00
Dave Mateer c772082f0e counter_screenshots to counter_warc_files in wacz_extractor so don't get error about add mulitple items with same id. 2025-06-03 12:34:41 +01:00
msramalho ee68f3efee
Merge remote-tracking branch 'origin/main' into feat/seleniumbase 2025-06-03 11:05:16 +01:00
dependabot[bot] efe2a1a8b6
Bump the actions group in /scripts/settings with 4 updates
Bumps the actions group in /scripts/settings with 4 updates: [@mui/icons-material](https://github.com/mui/material-ui/tree/HEAD/packages/mui-icons-material), [@mui/material](https://github.com/mui/material-ui/tree/HEAD/packages/mui-material), [react](https://github.com/facebook/react/tree/HEAD/packages/react) and [react-dom](https://github.com/facebook/react/tree/HEAD/packages/react-dom).


Updates `@mui/icons-material` from 6.4.12 to 7.1.1
- [Release notes](https://github.com/mui/material-ui/releases)
- [Changelog](https://github.com/mui/material-ui/blob/master/CHANGELOG.md)
- [Commits](https://github.com/mui/material-ui/commits/v7.1.1/packages/mui-icons-material)

Updates `@mui/material` from 6.4.12 to 7.1.1
- [Release notes](https://github.com/mui/material-ui/releases)
- [Changelog](https://github.com/mui/material-ui/blob/master/CHANGELOG.md)
- [Commits](https://github.com/mui/material-ui/commits/v7.1.1/packages/mui-material)

Updates `react` from 19.0.0 to 19.1.0
- [Release notes](https://github.com/facebook/react/releases)
- [Changelog](https://github.com/facebook/react/blob/main/CHANGELOG.md)
- [Commits](https://github.com/facebook/react/commits/v19.1.0/packages/react)

Updates `react-dom` from 19.0.0 to 19.1.0
- [Release notes](https://github.com/facebook/react/releases)
- [Changelog](https://github.com/facebook/react/blob/main/CHANGELOG.md)
- [Commits](https://github.com/facebook/react/commits/v19.1.0/packages/react-dom)

---
updated-dependencies:
- dependency-name: "@mui/icons-material"
  dependency-version: 7.1.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
- dependency-name: "@mui/material"
  dependency-version: 7.1.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
- dependency-name: react
  dependency-version: 19.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: actions
- dependency-name: react-dom
  dependency-version: 19.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: actions
...

Signed-off-by: dependabot[bot] <support@github.com>
2025-06-02 20:21:07 +00:00
Miguel Sozinho Ramalho 6735fa890b
v1.0.1 dependency updates, generic extractor improvements (#307)
* wacz: allow exceptional cases where more than one resource image is available

* improves generic extractor edge-cases and yt-dlp updates

* REMOVES vk_extractor until further notice

* bumps browsertrix in docker image

* npm version bump on scripts/settings

* poetry updates

* Changed log level on gsheet_feeder_db started from warning to info (#301)

* closes 305 and further fixes finding local downloads from uncommon ytdlp extractors

* use ffmpeg -bitexact to reduce duplicate content storing

* formatting

* adds yt-dlp curl-cffi

* version bump

* linting

---------

Co-authored-by: Dave Mateer <davemateer@gmail.com>
2025-06-02 20:57:12 +01:00
msramalho 69028588b3
linting 2025-06-02 20:04:34 +01:00
msramalho b351a33593
version bump 2025-06-02 20:03:48 +01:00
msramalho 87e1cdc102
adds yt-dlp curl-cffi 2025-06-02 20:02:35 +01:00
msramalho 4170c2011c
formatting 2025-06-02 19:33:55 +01:00
msramalho dd4e372703
use ffmpeg -bitexact to reduce duplicate content storing 2025-06-02 19:33:53 +01:00
msramalho b9f7927a3b
closes 305 and further fixes finding local downloads from uncommon ytdlp extractors 2025-06-02 19:14:09 +01:00
msramalho d99b7c9efe
Merge remote-tracking branch 'origin/main' into dev 2025-06-02 13:25:34 +01:00
Dave Mateer 48be13fb2a
catch for if self.comments are true but no actual comments in video (#303)
* catch for if self.comments are true but no actual comments in video

* simplifies check code

---------

Co-authored-by: Miguel Sozinho Ramalho <19508417+msramalho@users.noreply.github.com>
2025-06-02 13:02:19 +01:00
Dave Mateer 4aae5047f5
Changed log level on gsheet_feeder_db started from warning to info (#301) 2025-06-02 12:53:21 +01:00
msramalho 258e56aa26
poetry updates 2025-06-02 12:52:26 +01:00
msramalho 9ad6213efa
npm version bump on scripts/settings 2025-06-02 12:47:31 +01:00
msramalho 2f36e50e0b
bumps browsertrix in docker image 2025-06-02 12:06:14 +01:00
msramalho 2d7206f99d
REMOVES vk_extractor until further notice 2025-06-02 12:06:02 +01:00
msramalho ac24fd8f49
improves generic extractor edge-cases and yt-dlp updates 2025-06-02 12:03:51 +01:00
msramalho ee3e871dd8
wacz: allow exceptional cases where more than one resource image is available 2025-05-28 11:53:29 +01:00
msramalho e6fdef66df
improves instructions on docker setup with an example URL 2025-04-28 11:16:01 +01:00
msramalho 5cf640af8a
experiments with seleniumbase 2025-04-28 11:08:00 +01:00