Wykres commitów

27 Commity (main)

Autor SHA1 Wiadomość Data
R. Miles McCain f603400d0d
Add direct Atlos integration (#137)
* Add Atlos feeder

* Add Atlos db

* Add Atlos storage

* Fix Atlos storages

* Fix Atlos feeder

* Only include URLs in Atlos feeder once they're processed

* Remove print

* Add Atlos documentation to README

* Formatting fixes

* Don't archive existing material

* avoid KeyError in atlos_db

* version bump

---------

Co-authored-by: msramalho <19508417+msramalho@users.noreply.github.com>
2024-04-15 19:25:17 +01:00
Miguel Sozinho Ramalho 7a21ae96af
V0.9.0 - closes several open issues: new enrichers and bug fixes (#133)
* clean orchestrator code, add archiver cleanup logic

* improves documentation for database.py

* telethon archivers isolate sessions into copied files

* closes #127

* closes #125

* closes #84

* meta enricher applies to all media

* closes #61 adds subtitles and comments

* minor update

* minor fixes to yt-dlp subtitles and comments

* closes #17 but logic is imperfect.

* closes #85 ssl enhancer

* minimifies html, JS refactor for preview of certificates

* closes #91 adds freetsa timestamp authority

* version bump

* simplify download_url method

* skip ssl if nothing archived

* html preview improvements

* adds retrying lib

* manual download archiver improvements

* meta only runs when relevant data available

* new metadata convenience method

* html template improvements

* removes debug message

* does not close #91 yet, will need a few more certificate chaing logging

* adds verbosity config

* new instagram api archiver

* adds proxy support we

* adds proxy/end support and bug fix for yt-dlp

* proxy support for webdriver

* adds socks proxy to wacz_enricher

* refactor recursivity in inner media and display

* infinite recursive display

* foolproofing timestamping authortities

* version to 0.9.0

* minor fixes from code-review
2024-02-20 18:05:29 +00:00
Miguel Sozinho Ramalho 3e56ef137d
reduce s3 duplicating while keeping random urls via hash (#112) 2023-12-12 19:12:03 +00:00
Galen Reich 381940f5a8
Fix Selenium headless invokation (#106)
Co-authored-by: msramalho <19508417+msramalho@users.noreply.github.com>
2023-11-13 11:56:35 +01:00
msramalho ceb717ea65 exclude vk emojis 2023-08-17 18:11:26 +01:00
msramalho 6e4fb76940 exclude ok resource images from wacz enricher 2023-08-09 11:26:46 +01:00
msramalho 60a1f3a27a minor fixes 2023-07-31 16:08:48 +01:00
msramalho fb197f1064 excluding telegram embeds 2023-07-28 12:57:15 +01:00
msramalho aa71c85a98 improving ignored content from waczs 2023-07-28 12:19:14 +01:00
msramalho fc93ebaba0 cleanup 2023-07-28 10:49:39 +01:00
msramalho 59551b3b20 minor improvements: finding best twitter image quality 2023-07-27 21:36:15 +01:00
msramalho f086d89111 new escape message 2023-07-27 20:14:59 +01:00
msramalho dd034da844 feat: WACZ enricher can now be probed for media, and used as an archiver OR enricher 2023-07-27 15:42:10 +01:00
msramalho 888ad8f004 fix: twitter hack videos extension detection 2023-07-26 16:12:56 +01:00
msramalho 7c0b05b276 new column 2023-06-26 17:27:57 +01:00
msramalho 613b1f1e50 properly overwrite configs 2023-05-19 12:35:19 +01:00
msramalho a655b3c987 gsheet accepts ID too 2023-05-19 12:17:34 +01:00
msramalho 45b982ec38 fix: max chars on sheets cell 2023-05-10 18:57:33 +01:00
msramalho ae3e607705 fix: depreacating thumbnail_index 2023-05-09 11:29:05 +01:00
msramalho c1a60fde8a fix: deprecates duration column 2023-05-09 11:26:19 +01:00
msramalho 23894fad51 normalize columns 2023-02-20 16:08:35 +00:00
msramalho aa5430451e instagram archiver via telegram bot 2023-02-17 15:46:29 +00:00
msramalho 5505255ea3 url auth wall detect 2023-02-17 15:45:58 +00:00
msramalho 5b0593ce82 arg parse fix 2023-02-02 11:00:24 +00:00
msramalho d1e4dde3f6 fixing imports 2023-01-27 00:19:58 +00:00
msramalho 746f6a333e further cleanup 2023-01-21 19:57:54 +00:00
msramalho 753039240f pyproject 2023-01-21 19:01:02 +00:00