Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm not sure the research here is done. "perspicacious", a dictionary word, gets three x's; bj3UJ@P8uy9XD, a password of the same length I just generated with LastPass (and do not and will not use, of course), gets two (on the j and the 9). (In both cases, this was the first thing I tried of that length; I did not go hunting for a good dictionary word or hunting for a "bad" random password. If I got particularly (un)lucky, it's legitimately so.)

It's not really the fact that perspicacious is a dictionary word that I'm complaining about; I get that this is basing its decisions on a predictive algorithm and we can obviously augment this with some of the more traditional heuristics as well. What confuses me is that perspicacious, while perhaps being a $10 word, is also phonetically fairly normal, which ought to be bad. I would, for example, expect the s at the end to be fairly expected, or the a after the c.

I salute the idea, though. I've tried to make the point that a markov-esque password guesser could guess a great deal more passwords than we've seen so far. (Recall that while HN may frequently see Markov techniques used to randomly walk the nodes, you can also use it deterministically to enumerate possibilities in probability order quite easily.) It would be great to have something to simply point at.



It seems that rather than having some threshold of predictability, they are just taking the three most likely characters. So even though "ca" is a very common digraph, it isn't in the top 3 in the position you had it.

The difficulty, since this is intended a user aid, seems to be in balancing soundness (using a threshold) with user experience (sometimes predicting 3 letters, sometimes predicting 10, which would be overwhelming).




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: