Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Does this effectively mean Padmapper can use CL data once again?


CL has also removed themselves from search engines, so PadMapper can no longer use the google cache workaround.

I don't think they can go back to scraping either, but I'm not certain of that aspect of it.

This may be the real reason behind the change. Perhaps CL just doesn't need it any more.


That sucks, searching is how I double-checked dubious craigslist posts. Select a couple phrases and search google. If identical posts pop up under various other cities on craigslist, then I knew for sure it was spam/fraud.


Just one of the many ways third parties fix broken things on Craigslist. Things Craigslist refuses to recognize.


I think this is correct, though the EFF post is too vague, and doesn't quote the passage in its current state.

Also, CL has not removed itself from search engines. They're using NOARCHIVE.


Last I checked, robots.txt excludes all crawlers from at least several major classifications, including housing listings (/hhh):

    User-agent: *
    Disallow: /cgi-bin
    Disallow: /cgi-secure
    Disallow: /forums
    Disallow: /search
    Disallow: /res/
    Disallow: /post
    Disallow: /email.friend
    Disallow: /eaf
    Disallow: /reply
    Disallow: /?flagCode
    Disallow: /ccc
    Disallow: /hhh
    Disallow: /sss
    Disallow: /bbb
    Disallow: /ggg
    Disallow: /jjj
    Disallow: /*rss$
http://www.craigslist.org/robots.txt

'jjj' is jobs, 'ggg' is gigs, 'bbb' is services, 'sss' is for sale, 'ccc' is community, 'res' is resumes.



Point. I mistakenly assumed /aaa covered all of housing.

Looks like the same principle applies to a number of other subcategories. So while overview indices aren't spiderable, subindices are.


Google still provides a screenshot of CL pages - I doubt it would take much work for an enterprising hacker or 3taps to hook an OCR library up to that.


But wouldn't that be insane on bandwidth?


Inbound bandwidth is free.


I'm talking about for downloading all those pictures.


That's the point he was making - inbound traffic to AWS is $0/GB.


The exclusive license effectively meant that you couldn't cross-post to both CraigsList and PadMapper because you had granted exclusivity to CraigsList.


Ok makes sense, so that legitimizes something like PadLister. Does this however enable one to scrape CL for data?


...this does not legitimize anything in particular. It says it's OK for users to post ads on multiple sites, not that other sites can scrape CL's user content.

This is akin to freelance contract work: if you're a news photographer, there's usually a clause saying that while you can own your photos, you may not sell them to your client's competitors (at least right away).

So CL is saying to its users: Go ahead and copy and paste your stuff to other sites. So, not a huge difference by any means in what most users assumed.


PadMapper was doing the cross-posting to PadList to slowly build its user base. I however did not realize that Craigslist frowned upon that until the change in the agreement.


PadMapper doesn't take CL listings and put it into Padlister. What's listed on Padlister was listed by users.

PadMapper puts CL listings and Padlister listings on the same map.


My understanding is that this is a separate legal issue. CL's complaint against PadMapper is for scraping, not with users reposting on PadLister.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: