Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Huh, I I have been working on solution to that problem.

My project allows to define rules for various sites, so eventually everything is scraped correctly. For YouTube yet dlp is also used to augment results.

I can crawl using requests, selenium, Httpx and others. Response is via json so it easy to process.

The downside is that it may not be the fastest solution, and I have not tested it against proxies.

https://github.com/rumca-js/crawler-buddy



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: