Patrick Robertson
dba44b1ac1
Use WebDriverWait when waiting for elements in screenshot enricher
2025-03-07 12:07:54 +00:00
Patrick Robertson
e756f1504f
Remove geckodriver .tar file
2025-03-07 11:52:14 +00:00
Patrick Robertson
a47e18ef9a
Bump gecko driver to 0.36.0
2025-03-03 16:00:11 +00:00
Patrick Robertson
0dfab2d1bc
Add some code to attempt to click the cookies banners on various websites
2025-03-03 15:55:04 +00:00
Patrick Robertson
dea0a49600
Download correct gecko-driver for the platform + fix setting executable path when running in Docker
...
Fixes #232
2025-03-03 15:41:44 +00:00
Erin Clark
011ded2bde
Merge pull request #225 from bellingcat/small_issues
...
## GSheets Columns updates
- Update the available columns in the Google Sheet Feeder and Database.
- Update the Sheet Template to reflect this.
## Other Fixes
- Ensure test file cleanup.
- Additional tests.
- Correctly mark download test.
- Small typos.
2025-03-03 13:06:27 +00:00
erinhmclark
4280791f07
Fix mocking in test_wayback_enricher.py.
2025-02-27 11:25:58 +00:00
erinhmclark
8124bb831d
Merge branch 'main' into small_issues
...
# Conflicts:
# src/auto_archiver/core/base_module.py
# src/auto_archiver/utils/misc.py
2025-02-26 13:19:49 +00:00
erinhmclark
b2e654aef9
Remove context manager from test_pdq_hash_enricher.py
2025-02-26 12:57:33 +00:00
erinhmclark
9157846930
Add docstrings to explain date formats.
2025-02-26 10:01:52 +00:00
erinhmclark
35b5ab2eb1
Update poetry.lock
2025-02-25 20:17:48 +00:00
erinhmclark
83a08dd215
Update date parsing to use dateutil.parser in misc.py
2025-02-25 20:17:31 +00:00
erinhmclark
9bc6dd5c3c
Add set_content into generic_extractor.py.
2025-02-25 20:07:00 +00:00
erinhmclark
cf1219f798
Add text content into gsheet.
2025-02-25 20:06:44 +00:00
Patrick Robertson
1ad158c016
Merge pull request #211 from bellingcat/docs_improvements
...
Docs tidyups, howto on logging and authentication, remove exit(), small fixes
2025-02-25 14:13:13 +00:00
erinhmclark
1df5129268
Small typos.
2025-02-25 14:08:38 +00:00
erinhmclark
73b434aafc
Tests for test_vk_extractor.py.
2025-02-25 14:08:28 +00:00
erinhmclark
2d276cb9c4
Fix tmp test file.
2025-02-25 14:08:14 +00:00
Patrick Robertson
d10c7fbe55
Better documentation based on the discord feedbackgst
2025-02-24 22:42:57 +00:00
Patrick Robertson
ca1ed418aa
Throw an error for invalid __manifest__ syntax + fix: allow default values of False/None
2025-02-24 21:46:24 +00:00
Patrick Robertson
73a2e2d752
Fix tests for moving orchestration to secrets/orchestration.yaml
2025-02-21 19:05:39 +00:00
Patrick Robertson
091a19e25c
Further docs improvements/tidy ups
2025-02-21 16:52:30 +00:00
Patrick Robertson
77212e8e3f
Finishing touches to the how-tos
2025-02-20 15:45:48 +00:00
Patrick Robertson
9661e90a05
Allow disabling logging in auto_archiver with logging: enabled: false
2025-02-20 15:45:32 +00:00
Patrick Robertson
0bec71d203
Finish how to on authentication
2025-02-20 15:33:50 +00:00
Patrick Robertson
4174285898
Fix unit tests
2025-02-20 13:18:06 +00:00
Patrick Robertson
eda359a1ef
Fix json loader - it should go in 'validators' not 'utils'
...
Fixes #214
2025-02-20 13:10:39 +00:00
Patrick Robertson
40488e0869
Use 'Auto Archiver' naming for consistency.
...
auto-archiver is reserved in the docs for when talking about the command line usage
2025-02-20 11:50:29 +00:00
Patrick Robertson
061f29c885
How-to on updating config file to version 0.13+
2025-02-20 11:46:57 +00:00
Patrick Robertson
cbea551876
Better display name for wayback machine to emphasise it's typically used as an enricher
2025-02-20 11:46:57 +00:00
Patrick Robertson
b978484a89
Rename wacz_enricher to wacz_extractor_enricher. Fixes #205
2025-02-20 11:46:57 +00:00
Patrick Robertson
49b6c32058
Fix the 'full' mode which creates a complete config file
2025-02-20 11:34:05 +00:00
Patrick Robertson
4b51ec9ad5
Remove dangling import
2025-02-20 11:20:16 +00:00
Patrick Robertson
7734a551fa
Move 'assert_valid_url' out into utils, don't use assert but raise
...
assert is recommended only for debugging
2025-02-20 11:19:29 +00:00
Patrick Robertson
77b2b099c6
Replace exit() with raise exceptions. Better for code implementations
...
exit() is reserved solely for command line-called areas now
also assert is only recommended for debugging
2025-02-20 11:19:13 +00:00
Patrick Robertson
40b8359348
Implementation test with 2 x orchestrators with different configs
2025-02-20 11:18:28 +00:00
Patrick Robertson
5ccea8e44a
Absolute paths in README for Github/PyPi/Dockerhub etc.
2025-02-20 11:18:28 +00:00
Patrick Robertson
7dde8d609d
Merge main
2025-02-20 10:29:57 +00:00
Patrick Robertson
6ea943b680
Fix link
2025-02-20 10:27:24 +00:00
Patrick Robertson
5211c5de18
Merge pull request #210 from bellingcat/logger_fix
...
Fix issue #200 + Refactor _LAZY_LOADED_MODULES
2025-02-19 15:11:42 +00:00
Erin Clark
6cdefaa751
Merge pull request #194 from bellingcat/tests/add_module_tests
...
Add unit tests for individual modules.
Includes a couple of small bug fixes and light refactoring.
2025-02-19 13:51:43 +00:00
Patrick Robertson
04507577b6
Version bump
2025-02-19 13:36:50 +00:00
erinhmclark
47a634fc63
Add WACZ, Wayback and local storage tests.
2025-02-19 13:14:08 +00:00
Patrick Robertson
a9802dd004
Remove the global _LAZY_LOADED_MODULES and allow each instance of ArchivingOrchestrator to load its own modules
2025-02-19 12:25:35 +00:00
erinhmclark
a8ffb19325
Fix auth key name for cookies_from_browser.
2025-02-19 10:40:54 +00:00
Patrick Robertson
222a94563f
WIP: Docs tidyups+add howto on logging and authentication
...
(Authentication is WIP)
2025-02-19 10:37:04 +00:00
Patrick Robertson
eb60b271b9
Fix issue #200
2025-02-19 10:35:14 +00:00
erinhmclark
ddf2e76624
Include Atlos Storage __init__.py for module recognition.
2025-02-19 09:24:34 +00:00
erinhmclark
10a5ad62b8
Include Atlos tests, metadata fixture.
2025-02-19 09:18:41 +00:00
erinhmclark
f0fd9bf445
Updates tests to use pytest-mock.
2025-02-18 23:32:03 +00:00