Hacker Newsnew | past | comments | ask | show | jobs | submit | woadwarrior01's commentslogin

aka Normalized compression distance (NCD). Its close cousin: Normalized Google distance (NGD) is also super interesting!

https://en.wikipedia.org/wiki/Normalized_compression_distanc...


I have fond memories of using it for PIC16 as a teenager, ~25 years ago. It was a huge step up from writing assembly. My parents bought me a Microchip PICSTART Plus, but refused to pay for a Hitech-C compiler license. :)

Cool, I actually applied for a diploma thesis for doing the PIC 16 port around 2004. Raphael Neider actually was quicker then me, so he maintained the port for our lab at the Uni Karlsruhe (i am still in the same lab). We needed the port because was needed to allow people to compile our OS for our sensor nodes at the time. Still did a lot of testing of the PIC 16 port at the time because my thesis was about implementing a filesystem and a REST like API for the nodes.

That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.

https://x.com/EricSimons/status/2099252922098061714



Incidentally, there's a nascent OSS project called OpenSwiftUI.

https://github.com/OpenSwiftUIProject/OpenSwiftUI


If you care so much about user privacy, why do you exfiltrate events named "dictation_started", "dictation_completed", "dictation_rejected", "dictation_blocked", "insertion_failed", etc to PostHog, with properties named: "audio_seconds", "latency_ms", "duration_ms", "install_id", etc?

Because anyone building a system needs performance metrics to understand how it’s behaving in aggregate / if they update the model? With the exception of install_id these all seem innocuous. Install_id is maybe needed anyway for purchase management. Probably could use a session uuid to track the other metrics at the cost of blinding you to being able to find individual customers with a bad experience which is probably much more important for a small player like this.

Apple provides APIs for doing this in a privacy-sensitive way, instead of exfiltrating data like this, carte blanche using PostHog.

https://developer.apple.com/documentation/metrickit


Thanks for your Feedback. More Privacy Improvements, Package Corrections and Attribution Incoming.

Yes These metrics are there to check system issues. Like Crashes and etc. Also I need to check Model performance too. However your data is yours, and this data will be anonymized too. I really want Jexxa to be private and secure, as I am a heavy user of it too.

I spent ~10 minutes looking at it.

Vibe-coded website.

Both bundles contain quantized versions of the CohereLabs/cohere-transcribe-03-2026 model. The model in JEXXA.dmg is MLX int8 g64 affine quantized. JEXXA-Small.dmg has the same model but MLX int4 g32 affine quantized. There's also an int8 quantized WeSpeaker ECAPA-TDNN speaker-verification ONNX encoder. No acknowledgments for both models (doesn't the Apache 2.0 license require it? WeSpeaker's cc-by-4.0 license certainly does). Both bundles also contain full-blown Python 3.12.8 runtimes with about a dozen packages installed.

Apps are unsandboxed menubar apps and also contain PostHog analytics and Supabase auth. So, I wouldn't run it on any of my machines.

A few engineering hygiene issues like a .pytest_cache directory, Python code, .DS_Store files, etc.

Given the above, I suspect the "I am training better models." is just marketing speak.

nb: The demo link without LinkedIn tracking slop: https://www.linkedin.com/posts/sankyde_jexxa-demo-httpsjexxa...


Very interesting!

I suppose the next logical thing to do would be to train a language model to generate random flags, call it Artificial Flag Intelligence (AFI) and raise a $10m pre-seed at $100m post. :P


I'm not defending claudeslop, but TBF, anaphora and its lesser-known sibling: epistrophe far predate the English language, let alone something new like LLMs.

"I'm not defending claudeslop, but <defends slop>". People obviously know this. The problem is that using a dramatic rhetorical flourish every third sentence for completely un-dramatic things is fucking tiring. And so are the people defending this shit. "Oh, humans use em-dashes too, so actually there's literally no difference between a human using one or two and Claude using 400 em-dashes in one article". A human using things correctly and a language model using them incorrectly are not the same fucking thing.

> The problem is that using a dramatic rhetorical flourish every third sentence for completely un-dramatic things is fucking tiring.

100% this. It's exhausting reading this stuff, and those dramatic flourishes are a big part of why.


I believe that what the previous commment is trying to say is that even when we identify phrases highly attributable to llm output we can never be certain. Which is the nefarious thing about broad use of computer generated text. I noticed comments here on hacker news containing conspicous mistakes which llms hardly ever do anymore and was wondering whether it was people purposefully adding mistakes to identify as noslop or whether llms got worse or if it was bots trying to fool me. My misspelling of conspicuous was left in on purpose in this case. All i can say is "trust me bro, I am not ai".

The issue is the frequently repeated, formulaic use.

Repetition at the start or end (i.e. Anaphora/Epistrophe) are not used lightly in prose or verse, they serve a rhetorical purpose - usually add rhythm or strengthen the theme.

The type of repetition Claude uses is what Fowler[1] called "elegant variations" and discouraged in modern style guides for a good reason - people find it really annoying.

[1] Henry not Martin


On a related note, I came across this USB-C KVM on pre-order the other day. It ostensibly supports Tailscale, which would be really useful.

https://sipeed.com/nanokvm-go


I own 3 of these for my homelab. They work really well w/ Tailscale! They sell some PCIE ones, too.

Was posting "Now make it join Tailnets." to the comments then saw your remark.

TY for link.


I have the predecessor at home, it's solid and the tailscale support is great!

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: