Hacker Newsnew | past | comments | ask | show | jobs | submit | jdefr89's commentslogin

I implemented it quite quickly I use SimHash for now but I need to refactor and possibly replace…


Wow. Wasn’t expecting to see this here… I am the author…


Quick question, did you use an LLM to help with this? (Not judging one way or another).

The only reason I ask is that for the past couple of weeks I’ve been making a similar utility to test out the Sol model and it’s hilarious how similar the architecture and even the CLI switches are to what Codex came up with.


(Sorry if I missed anything obvious from the documentation)

Does this work on Windows


You can still exploit a system and easily prove it via simply popping a shell or calc.exe or updating a database with a new entry, etc… They didn’t have to let it loose on the network. If that system was air gapped - problem solved.


But that’s the problem with AI it is like 16yo script kiddy who will just exfiltrate all your PII and think it did good job. Mature pentester would pop calc.exe make screenshot and be done.

Other problem is setting up air gapped test environment is a lot of work, especially if you expect it to be equal to real thing.

This pentest with AI is not as useful if you set up a single app - it really is useful if you want to find exploitable chains of exploits that seemingly might not be exploitable separately or not leading to full hack separately.


Security Researcher here. While you’re correct that air gaps aren’t a totally secure mechanism to rely on, they sure as hell can raise the bar for realistic exploitation. You pretty much need to rely on tricking someone into running your exploit or something of that nature. That said, they could have completely avoided this problem with an air gap. Simply don’t provide it network access. That isn’t too hard to do.


> That isn’t too hard to do.

No it’s not, and they aren’t fucking stupid. They were obviously courting this possibility so they could have another big headline. And it just happened to attack HF? I’d honestly be astonished if it wasn’t entirely deliberate.


I am at a similar salary. I live in Kendal Square; nearly on MIt campus for convenience and due to the fact I relocated to Cambridge from PA when I took my research pos at MIT lab. I pay 4000k a month for a STUDIO. That includes no utilities or anything. Then add groceries and all that stuff. Then add supporting your significant other and various other things…. trust me it disappears fast. Yes there is some cash left over if you do absolutely nothing but stay in your home… Things get tight. I shouldn’t feel pressure given my salary but sometimes I think if things really went downhill… my salary isn’t leaving me room for a care/worry free life style. Not the way it would have 10 to 15 years ago. Things are too expensive…


Hello Joseph. I am awesome Joseph. I am at MIT LL. Can’t wait to check this out a bit deeper.


A few things here.

1. Use a proper Markdown parser. The grammar is easy to define EBNF style most implementations I see now days use some recursive descent parser, etc… Regex implementation was used in original authors parser when it became popular.

2. You can resolve ambiguities and define more consistent symbols that make sense. Most markdown implementations are decent and follow common sense best practice syntax.

3. The beauty is its simplicity. You can pick it up in a few minutes and utilize it anywhere you damn near see a text box.

4. Parsing to HTML isn’t the only option! I mostly use TUI markdown viewers that render the document via ANSI escape codes beautifully. Check out glow project. But once you parse and have a AST, you can easily walk it and render it in other ways as well. Again though. Everyone can read a text document and a html document. You can render it to a PDF if need be.

5. Do we really need a whole new markup text2<format of some kind>? Markdown is simple fast and widely supported. So I have to say.. I prefer it over most things and that includes Rst.

If you need real beauty and power you can move to LaTeX or something… My two cents anyway.


I recently tried to create markdown like parser, I had to refer a lot to common marks. What I saw was madness. I never knew from casual use, that markdown is so complex. There is literally zero thought about parsing, it forces natural looking text into a format that can be structured into actual markup, but it has so many pitfalls and edge cases that it just feels wrong. Each time I've looked up another markdown parser I was stunned how overegineered they are. If you need an AST to parse something seemingly simple like markdown, then something went wrong. Probably every coreutil is smaller then the average markdown parser. Why?


> If you need an AST to parse something seemingly simple like markdown, then something went wrong.

This may surprise you, but that is very common in programming languages.


Doesn't surprise me at all.


Because Markdown is made for humans, not for machines.


Because the size of the optional parser for Markdown is irrelevant to the purpose of Markdown.


Sort of hard to do because AI it shoved down your throat in one form or another virtually everywhere you go. I also think a lot of us Hackers are mourning the fact we spent many years mastering machines and programming just to have the skill devalued (at least from the publics perspective) nearly over night. I personally think it is more important now more than ever to understand technology. To be able to write code, understand how a CPU works etc. Tech literacy will help prevent doom scenarios. A future where virtually everyone depends on AI and Computers but lacks people who actually understand them from a low level perspective seems bleak. I know thinking itself seems to have gone out of fashion and its given rise to misinformation and/or political nonsense like the rise of fascism etc... I think a lot of us just feel "empty" and are trying to express it.


I get it. I’ve been doing this for 11 years. I use agents everyday at work now and deal with all the benefits and problems of that. The craft is certainly changing and it will take years for everything to shake out and settle. I understand the desire to publicly wax poetic, but nobody actually knows shit about where we will land, so it gets a bit tiresome to see over and over.


I agree that humans should continue to value various forms of literacy even in the face of AIs that can do everything better than us. I too will continue to dig deeper into tech literacy. There was a Terence Tao paper recently that mentioned we are in a shift similar to the end of heliocentrism. It made clear that Earth is not the center of the universe, but Earth is still deeply valuable and important for humans. Much the same way that AI may supersede our understanding and intellect and make the are limitations more apparent, but our human intellect is still important to humans. Plus, what are you going to do when the price of LLM tokens are through the roof or you get messages like "burn an extra 1,000,000 tokens for a better implementation!".


I have some amount of hope that local open models with sufficient quantization are the future as hardware becomes more powerful and models become more optimized. I don’t think we will be living in thin client land forever. Human expertise and intelligence will continue to be important and anyone who says otherwise is being disingenuous.


Agreed. I am crossing my fingers that local open models can catch up in the future. Otherwise the big LLM companies will have everyone by the balls.


Over reliance on LLMs is going to become such a disaster in a way no one would have thought possible. Not sure exactly what, who, when, or where.. Just that having your entire product or repo dependent on a single entity is going to lead to some bad times…


> on a single entity

Contrary to the popular opinion here, there are other services beyond Claude Code. These usage limits might even prompt (har har) people to notice that Gemini is cheaper and often better.


On-premise LLMs are also getting better and likely won’t stop; as costs go up with the technical improvements, I would imagine cost saving methods to also improve


I still think it's basically unavoidable that most people who might pay for api access will end up on-prem.

Fixed costs, exact model pinning, outage resistant, enshittification resistant, better security, better privacy, etc...

There are just so many compelling reasons to be on-prem instead of dependent on a 3rd party hoovering up all your data and prompts and selling you overpriced tokens (which eventually they MUST be, because these companies have to make a profit at some point).

If the only counterbalance is "well the api is cheaper than buying my own hardware"...

That's a short term problem. Hardware costs are going to drop over time, and capabilities are going to continue improving. It's already pretty insane how good of a model I can run on two old RTX-3090s locally.

Is it as good as modern claude? No. Is it as good as claude was 18 months ago? Yes.

Give it a decade to see companies really push into the "diminishing returns" of scaling and new models... combined with new hardware built with these workloads in mind... and I think on-prem is the pretty clear winner.


These big players don’t have as big of a moat as they like to advertise, but as long as VC wants to subsidize my agents, I’ll keep paying for the $20 plan until they inevitably cut it off


gemini-cli has not been useable for weeks. The API endpoint it uses for subscription users is so heavily rate-limited that the CLI is non-functional. There are many reports of this issue on Github. [1]

1/ https://github.com/google-gemini/gemini-cli/issues?q=is%3Ais...


I use Gemini-CLI at work, and haven't noticed anything. I use Google Jules (free tier) on a toy project much more heavily and can't complain. I think sometimes the prompts take longer than they used to, but I couldn't care less. I'm not in a hurry.



Gemini better? What are y’all doing that it doesn’t crash and burn within the first minute of using it?

It might be acceptable for some general tasks, but I haven’t EVER seen it perform well on non trivial programming tasks.


Last time I used Gemini I watched it burn tokens at three times the rate of any other models arguing with itself and it rarely produced a result. This was around Christmas or shortly after.

Has that BS stopped?


It's still not uncommon for it to escape it's thinking block accidentally and be unable to end it's response, or for it to call the same tool repeatedly. I've watched it burn 50 million tokens in a loop before killing the chat.


No. It's still shit. It can do some well contained tasks, but it is very less usable on production codebases than gpt or claude models. Mainly because of the usage limits and the lack of good environments for us to use it on. Anthropic gets away with this because claude code, as bad as it is, is still quite functional. Gemini cli and antigravity are utter trash in comparison.


Exactly my experience. I remember thinking to myself that if this is what people get exposed to when they try to use a coding agent, no wonder there's so much bad-mouthing going on about LLMs. You use CC and you get usable output without much hassle, and in the end it costs way less because you aren't fighting with a substandard model.

Frankly, Gemini seems like Codex was two years ago. Lots of back and forth and nothing of value in the end.


For a second I hoped you were gonna comment on how LLMs are going to rot out our skillset and our brains. Like some people already complaining they "have to think" when ChatGPT or Claude or Grok is down.

Oh well.


The other day I was doing some programming without an LSP, and I felt lost without it. I was very familiar with the APIs I was using, but I couldn't remember the method names off the top of my head, so I had to reference docs extensively. I am reliant on LSP-powered tab completions to be productive, and my "memorizing API methods" skill has atrophied. But I'm not worried about this having some kind of impact on my brain health because not having to memorize API methods leaves more room for other things.

It's possible some people offload too much to LLMs but personally, my brain is still doing a lot of work even when I'm "vibecoding".


Ironically this is one of my main use cases for LLMs

“Can you give me an example of how to read a video file using the Win32 API like it’s 2004?” - me trying to diagnose a windows game crashing under wine


Exactly. I feel this is the strongest use case. I can get personalized digests of documentation for exactly what I'm building.

On the other hand, there's people that generate tokens to feed into a token generator that generates tokens which feeds its tokens to two other token generators which both use the tokens to generate two different categories of tokens for different tasks so that their tokens can be used by a "manager" token generator which generates tokens to...

And so on. It's all so absurd.


Unsurprising people complain.

"Thinking is the hardest work there is, which is why so few people do it" — attrib Henry Ford

Now we have tools that can appear to automate your thinking for you. (They don't really think, but they do appear to, so...)


“Thinking is to humans as swimming is to cats. They can do it, but they prefer not to.” - Kahneman


I read that as implied.


AI will totally rot our brains, just like television, video games, and the internet all did before.


Do you feel that television, video games and the internet had a negligible impact on our culture?


This but unironically.


I don't get this pov, maybe b/c I'm not a heavy Claude Code user, just a dabbler. Any LLM tool that can selectively use part of a code base as part of the input prompt will be useful as an augmentation tool.

Note the word "any." Like cloud services there will be unique aspects of a tool, but just like cloud svc there is a shared basic value proposition allows for migration from one to another and competition among them. If Gemini or OpenAI or Ollama running locally becomes a better choice, I'll switch without a care.

Subscription sprawl is likely the more pressing issue (just remembered I should stop my GH CoPilot subscription since switching to Claude).


There's so many different models, from hosted to local and there's almost no switching cost as most of them are even api compatible or supported by one of the gateways (Bifrost, LiteLLM,...).

There's many things to worry about but which LLM provider you choose doesn't really lock you in right now.


It should be abundantly clear that depending on a single entity will screw you royally, but obviously we don't learn from the mistakes of others. We are condemned to repeat history because we don't know it.


So, like, GitHub then?


Or Cloudfare or AWS


How can automatic slop-prevention be a disaster? It's a feature.


if you rely on the black box of bullshit... you deserve your own fate.


Which one was yours???


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: