Hacker Newsnew | past | comments | ask | show | jobs | submit | girvo's commentslogin

We can’t use fable at work, opus and Astra are as good as it gets.

It supports ACP, so you can use it with any ACP speaking harness including many that would work with local agents

My only gripe with them is how they locked down their m365 scooters, made repairing it a pain.

It’s called “HBF”, high bandwidth flash, and it’s on its way!

You target the US companies: if they can't use these Chinese models, then they're less of a danger for a now captive audience in the US (and the West generally).

This is already kind of the case: the big enterprises don't really want to touch the latest Chinese models. It's a real pain, personally, I want to use them at work!


1. China is a bigger market than the US for Ai, they are on pace to process 100Q tokens this year, roughly the same or more than the US big companies

2. Enterprise trends are towards open weights, several routers and vendors now have more than half the volume going towards open weights


Yes, but thats not something a company engaged in regulatory capture for themselves care about: especially if they're worried they'll be outpaced and overtaken by the Chinese labs. Which they will be, IMO.

they care because they know it unlikely open weights will be banned, and thus available to American companies, with regulatory capture (onerous requirements) being a "good enough" "ban" that their big models don't face real competition, regardless of the open weight origin. American companies make open weights too, they are equally threatening to Big Ai financials.

That will then create incentives for companies that consume AI tokens to counter lobby against those regulations.

Incentives exist, they are already lobbying and making counter public statement, like Jensen Huang of Nvidia.

His first tweet ever, from this last July

https://images.nvidia.com/pdf/Open-Weights-and-American-AI-L...


Show them you can burn tokens in seven sessions day and night with comparable results to Opus with less energy and less than 10 dollars a day, per dev.

We have. Unfortunately there are political realities that get in the way, and Bedrock for example doesn't have GLM 5.3 (Flash or otherwise) or anything new/useful

I do imagine it'll change, but it hasn't yet.


If it's hosted, all they know is "data goes to China".

Until profitable, reputable third parties host open models in the US with ZDR or they become plug-and-play for self-hosting at a modest cost, paying the US models is as much about data protection and liability as performance.


It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch

Because you can run Qwen 3.8 Flash Next, Laguna S 2.1 and other medium-sized models that simply don't fit on a 5090?

The follow on question is if it's making us all so much more productive, where is the increased revenue? As far as I can tell, it's mostly the AI labs seeing that, not everyone using them (modulo small founders building new things and doing okay, I think)

Personally I believe that this boost in productivity will not necessarily lead to greater revenue.

All the companies have the same access to AI, and AI is making them all better at doing what they were doing before (writing software). So some companies that use AI really well may be able to take market share from other companies that are slow. But I think this might just lead to a more intense competition for customers.

Kinda like what happened to music after digital recording became the thing - we have more music than before and the music is better, but being a musician became a much more intense competition to find an audience.


People won't want to hear this, but for FAANG things are probably similar. I'm not very convinced that Meta, for example, is actually creating new revenue. It's just consolidating a lot of the existing global ad spend, it's just wrecked newspaper classifieds and the like.

probably a naive question but then would we expect companies not using it to lose market share to competitors that are using it?

Well I don't really know what will happen in the future. My first thought is that what will happen depends on the company and industry. I can think of some non tech companies where AI might not matter for a while, like a car wash or a lumber mill. But long term likely yes for tech companies and companies where tech makes a difference.

If every software company saw a roughly equal improvement in productivity, they'd still be splitting the same customer base amongst each other - not clear whether revenue would actually go up for anyone.

Perhaps some new markets could be entered that weren't feasible before?


But there have been trillions invested in infra for AI. Surely those investors are going to need to see a return at some point?

apparently those trillions were necessary just to keep the boat staying afloat.

Isn’t cloudflare a pretty clear answer to these? They are launching more products than ever, some of them clearly vibe coded, and their revenue is exploding.

I think you could make that argument, yeah.

https://au.finance.yahoo.com/quote/NET/financials/

Interesting to look at, a decent example for sure.


Is that because more people are signing up to use Cloudflare as a MITM to block the increasing avalanche of AI-generated traffic though??

I don’t know exactly where the money comes from but they are clearly launching vibecoded products https://try.cloudflare.com/ and seem to be quite good at it. So they are shipping more for sure and it seems to help win them business/ expand their existing accounts.

It may not have to do with increased revenue, but it does have everything to do with reduced cost, especially developer's cost.

I bet every company is finding up how to level up their employees via AI, so that they can use less of them in the future.

So even without increased revenue, AI has its (mis)uses.


The increased revenue is in startups, like the recent couple YC batches.

Do you have a source for that? Are they raising more money or receiving more money for AI-oriented products or are they actually making more money on consumer/B2B end products?

And yet this linked page is filled with AI slop tells, annoyingly, so it makes it hard to separate the wheat from the chaff in terms of useful information.

Check out eugr’s TP=1 sparkrun recipe :)

It’s an NVFP4 quant, but it fits, and is surprisingly capable.


do you have a HF link? HF search is not uncovering it for me

(or is it somewhere else)


https://github.com/spark-arena/eugr-recipes/blob/main/recipe...

This one!

I'd recommend pointing your agent at it (after installing sparkrun), and asking it to research the absolute latest in TP=1 Flash-Next - mine grabbed particular vLLM nightlies and mods to improve performance, and it was well worth it.


I have a quirky vLLM on k8s on 2x OEM sparks setup with 9 models available to me. I'm not keen to run nightly vLLM, too many issues with it in the past. Going the qwen-next path means displacing things I use daily :/

I have a watchful eye on the diffusion ~ Jev/Kev PR

https://github.com/vllm-project/vllm/pull/57250


For what it's worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don't think you'll need the nightly for that at all, just use the recipe)

I'm so tempted to buy a second one...


prices have gone up quite a bit...

I'm running embedding, reranking, and policy tuned models too, and a Jev/Kev when that's landed. Flash Next is not a substitute for those

I have OpenCode/Fireworks to access big models



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: