Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sites like these are, in my opinion, the scourge of the internet. There is a lot of talk nowadays about curated search engines displacing machine-generated search engines but I tend to think this goes too far. A search engine that could reliably determine the authoritative source of duplicated content and only include that source would be killer. Seems within the realm of possible... Anyone working on that?


Look up information cascades.

We're doing something along a similar vein for LazyReadr. The idea is to merge news about the same story together, an important step from there will be deciding which is the authoritative source to display. Going from that to effective search isn't a large leap.


Thanks for the tip. My search turned up this resource, which looks awesome:

http://news.ycombinator.com/item?id=1986198


I suspect Google isn't doing this for legal reasons. They probably have to take care not to look discriminatory in any way.


It shouldn't be difficult to detect. Whichever site receives content first is likely the definitive source.


Also, in many cases (eg. this one) the duplicated content even links to the original source, so even simple pagerank should put stackoverflow significantly higher than its duplicates.


see http://webmasters.stackexchange.com/questions/6556/does-the-... for the primary outcome of this investigation




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: