Hacker Newsnew | past | comments | ask | show | jobs | submit | febed's commentslogin

Opera browser has this feature baked in a long time ago, before the Chromium era.

8086 and up! Amazing


The link is broken


Is the model response quality identical to the memory unconstrained model?


Yeah, it must be exactly the same. The same weights are used, nothing skipped or pruned. But it might have differ to MLX for greedy decode because of small floating-point nums difference


Which model do you think currently hits the highest size-to-performance ratio for agentic tasks like tool calling?


OK so I don't have enormous agentic coding experience yet (I'm still learning about the tech more than coding with it) but:

- The consensus is that if you have a properly kitted out PC with enough VRAM, the Qwen 3.6 27B dense model is the one to beat.

Bit slow on my M1 Max so I haven't bothered with it much but I have no reason to doubt the consensus. Prism ML's new Ternary Bonsai variant of it makes it much easier to play with this model in limited RAM, but in my own toy experiments I have seen Ternary Bonsai get very stuck in thinking loops. There is another post-train variant from BottleCap called ThinkingCap, which you could try.

- The Qwen 35B MoE model is really impressive for code generation. I personally would pick this one for an older Mac or a machine with smaller VRAM; it's pretty fast, has good built-in MTP, great tool-calling.

- The Gemma 4 26B MoE has a similar capability but biased more to writing than coding. I think it writes really well, and it seems to have good knowledge of e.g. WordPress coding, SQL etc. Tool-calling let it down for agentic coding, and I haven't retested it since they fixed that

- The dense 31B Gemma 4 is large and runs slowly on my machine, but has very good general knowledge, writes well, so it should I think be better than the Qwen 27B for research tasks, and it should now be pretty solid at tool-calling.

- If you don't have much VRAM, you are not doing much coding (e.g. you want short snippets) and you want to experiment with local LLMs and perhaps in particular image analysis, the Gemma 4 12B is fun. It has an integrated vision decoder which is very impressive. Codewise, it's going to fail on long context tasks.


Thanks a lot for the recommendations!


No worries. YMMV for coding things but I've learned a lot more about LLMs this way than I think I could have from just reading about and using cloud LLMs.


Can the same be done with qwen3.6-35b-a3b?


Yeah, the same ideas should work for qwen. You can try porting this engine to use Owen.

Owen 3.6-35b-a3b was my initial idea, but I switched to Gemma because of its simpler architecture and kernels


Curious if the same idea could work with gpt-oss-120b? So one could run at least slowly on a Mac


Yeah, gpt-oss-120b is also MoE, so the same ssd-streaming and caching ideas should work. Feel free to fork and try implementing it!


What exactly do you mean by data cleaning


The coolest project I’ve got this running on is improving the depicts metadata for photos on Wikipedia. A lot of times they won’t have the landmarks tagged correctly in a photo. So I will load in all the metadata that exists from each photo and the pixels of those photos and give a small qwen agent access to Wikipedia search as well as a geocoder. It does a great job of figuring out what is depicted and tagging it with the correct depicts field. Im still early on but I have been able to double the number of places that have a photo attached to them on wikimedia


Interesting project. So I guess the search tool is to crawl Wikipedia for articles? And how do you ensure that the tagging stays within the Wikidata taxonomy? How exactly are you using a geocoder? Sorry, just curious


So like oftentimes the picture will be of a church and there’s geographic coordinates for where the photo was taken. My qwen will use the geocoder to search for “church” at the coordinates of the photo and then read the Wikipedia articles about all the churches nearby and see if any of them could plausibly be the church. So far I have parsed about 2 million photos and have tagged about 800k places. My goal is to do the whole 40 million places to create a world map of open places with photos. The tool I’m using is topoloop for the geocoding


Would be interesting to make a similar one for Tibetan chants.


What SDK are you using? Or is it custom?


Off topic, but apart from pi-tui, is there a recommended TUI library that integrates well with the Pi? I want to have a multi pane TUI experience like lazydocker right inside Pi. Pi-tui is a bit limited.


Why not just have multiple instances of pi inside one of the many multiplexers out there? Tmux, Zelliij, herdr, for example.


No.

Use zelilij


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: