Hacker Newsnew | past | comments | ask | show | jobs | submit | myrmidon's commentslogin

It's an arab loan word for ~"price list" (originally strongly suggesting an import-duty-like meaning).

The same word exists in e.g. German or Italian, but mostly just means "pricing scheme" there, like in the article ("import duty" using a different expression in those languages).

Might be a mostly/somewhat "false friend" for a German (pv magazine is german).


Tarif/Tarife means price list of anything. Your barber, restaurant, taxi service, hospital can have one. tariff as in the meaning of custom tax is a very narrow meaning considering the usage of the word in general (but that's OK).

Import duty is "gümrük vergisi", and the general list of this is "gümrük vergisi tarifesi" in broad sense. It also has a more technical name here but both I don't remember it, and it doesn't matter in this context.

In the Japan's context we would call it "elektrik fiyat tarifesi", or "elektrik tarifesi" for short, for example.

Cheers,

Your friendly Turkish HN user.


Buying breaking bad on DVD is <3 months of netflix subscription (no ads).

So if you watch <2 seasons per month on average, you already come out ahead, and this is assuming only a single subscription.

People tend to wastly overpay for subscriptions because they end up underutilized and are typically never cancelled quickly enough because people are lazy.


Sadly there is no realistic approach to get there (even assuming fantasy-levels of technology).

Chemical rockets are currently the only thing that allow sustaining such accelerations briefly for human-size payloads, but the low exhaust velocity and exponential reaction mass requirement make sustaining it for days/months/years completely impossible.

You'd have to supply the energy externally (i.e. some sort of beam propulsion), but getting any significant fraction of g out of such a system (with human-sized payloads) seems unlikely within the next centuries, especially as the distance increases.


> You'd have to supply the energy externally (i.e. some sort of beam propulsion), but getting any significant fraction of g out of such a system (with human-sized payloads) seems unlikely within the next centuries, especially as the distance increases.

Or some form of ram scope, plenty of H everywhere, but those also have the issue of you collecting things while going at relativistic speeds.


You need ablative ships cruising ahead of you..

Just replace the model with a human student.

"Training" on textbooks => fine

"Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.


Presumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO).

The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work.

Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism.

Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".


To be clear, the accusation is that they trained on the chats they used while working on the problem. Not published work or even a preprint.

Your post does not distinguish, and it matters.


How does it matter? It either is or is not plagiarism. Ripping off a published textbook isn't somehow better than ripping off private correspondence. Both are serious acts of academic misconduct on account of the part where you knowingly and intentionally portrayed someone else's work as your own.

Note that I am not taking a stance on what openai allegedly did or did not do one way or the other. I am merely pointing out what I see as a fatal flaw in the line of argument presented by the earlier commenter - the idea that training on an item is on its own sufficient to establish plagiarism of it.


voluntarily is stretching it for an opt-out

This is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.

What about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?

Many do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here:

Science papers of a phd level must contain:

1. one or more novel insights

2. a long list of citations to contextualize them and

3. some work to prove that the insights are in fact meaningful

---

In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image.

That image is twice plagiarized:

1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit

2. the model failed to cite where it pulled the "horse" and "space" concepts from.

It merely did the work (3) to combine the concepts using the user provided insight.

---

The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits.

This is still academic plagiarism, even if you disagree that all LLM outputs are.


I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context.

As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus.

So is there any actual evidence that openai trained on the data in question? And further, did the openai proof directly build on someone else's novel insights as opposed to deriving everything from scratch? (I don't pretend to know but the vast majority of what I've seen so far in the comments here is what I'd characterize as brain-dead screeching. Certainly not the level of discussion I come to HN for.)

Separately, consider the implications of what you're arguing for there. Suppose your horse in a space suit picture were somehow valuable to society. Suppose that due to shortcomings of your tool you lacked the ability to readily and accurately identify the originators of the relevant concepts. Should you refrain from publishing this useful work due to the lack of citations? How are you supposed to handle this situation?

Remember that in this analogy everyone throughout society is on the same page that your tool consistently recycles other people's ideas while being technically incapable of producing reliable citations. The question is a simple trolley-esque problem - do you publish without proper citations for everyone's benefit and if so what are you supposed to say?


evidence that openai trained on the data: they would have denied it if they didn't train on it.

did the proof build on the insights:

the influence of an individual text in the training data is deeply weighted by quality, relevance, etc. a high quality proof in advanced mathematics written by a codex user is going to get boosted to the max.

the model is post-trained on prompt material. that is again going to boost it.

the prompt will boost this material specifically. perhaps they even rammed dense maths in particular into the model in post training.

anecdotally i have been able to get near-verbatim copies of original material out of models at inference. the type of work that buckmaster and alpoge fed into openai feels like the exact type of concept that would cause an "aha!" or "but what if?" in chain of thought. in fact i would bet that their work is in the logs.

the likes of astra and fable are thought to be up to 10T parameters in size. i consider it highly plausible that a semantic representation of the euler proof could be pulled out of the model weights in good shape.


I am not engaging further. You asked for meaningful refutation of your observation, I provided one.

Academia has stricter rules than regular society.


The chats the professor had are not generic knowledge. And yes of course you still need to cite Newton and Leibniz depending on what result you want to mention. What’s allowed to be not cited are not status quo, the term you’re looking for is “folklore” results aka results that have been around so long that 1) nobody knows who came up with them or 2) everyone knows who came up with them.

The second point: if you say you can’t prove that OpenAI actually used it, it doesn’t mean that OpenAI did not use it. It’s hacker news not lawyers news here lol. And OpenAI can’t prove that they didn’t use it either. The whole point is that Levent felt he had reasonable suspicion to believe the AI did use the result, because he felt like without his input on an unpublished paper it was unlikely for AI to reach the same result. I haven’t read the paper so I don’t know where I stand on that.

On the last point, about your “for the greater good” argument. It’s higher maths lol. I don’t know about this field but I doubt it’ll be very useful for society. Maybe it’ll make one part 2x faster which makes some rocket cheaper to launch. Does the average person care? Debatable. I think it’s reasonable to hold published papers in proof based fields to a higher standard. Otherwise the current & future problems of ML engineer fields just expand to other fields. No thanks.

Finally, if you anonpost to the autistic Internet forum that everyone else is “brain dead screeching”, it really just says something about yourself lol.


Even if AI used the result, AI pushed it to the finish line while Levent and Tristan did not. But I understand the approach was different, the information leak was only that it was "doable".

The approach was the same.

That's exactly the argument of the people calling it plagiarism machines. No-one ever really did refute it there was just a bunch of settlements for elite institutions so they weren't left empty handed like the various small time creators/authors etc were.

I think the bigger issue here is this feels like some PR smoothing happening that after all the work that went into "it's safe to use for enterprises" now we have what looks like openAI using private user data to scoop novel research and the question of why couldn't they do it for an enterprise with much more money on the line.


> now we have what looks like openAI using private user data to scoop novel research

Is there any actual evidence of that? All I've seen so far are empty accusations because "it would be in their interests" or whatever. Personally I'm inclined to believe that they honor their terms until it's demonstrated otherwise.


their terms grant them an irrevocable license to your data unless you specifically opt out of it.

What does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it

Depending on the education/age of new recruits, you might not have seen much of the, younger, lower scored cohort, because the sharp performance drop only really started after 2018 (so the oldest affected people are around 22).

You're right. I suspect there are class factors at play too. A lot of education happens in the home when you're privileged.

Modems becoming super cheap is gonna be a big problem for this, because those make it even harder to prevent such snooping.

Just imagine having to jam the devices in your own home just to stop them from profiling you.

Better pray to your deer god that modem + data contract is gonna stay more expensive that the expected value from selling more data...

We really need regulation to stop this.



Disconnect the u.fl to the antenna and it’ll be hard for it to transmit data unless you live next door to a cell tower.

It could include SMD antenna or PCB antenna.

I see the same thing with advertising. The cost of a banner ad in real life is always decreasing. Now every road intersection, shop aisle, plane seat etc has a huge super bright billboard blasting ads at people who can't realistically avoid it.

My house already jams cell signals at least

There is a comparison with "zstd -19" on the Silesia corpus, showing better compression ratio for bzip3 (47.2 vs 53MB) while being ~5 times faster (and using only half the memory).

Even if the examples are highly cherry-picked, it is quite suprising to me that such pareto-dominance is possible at all.

edit: Tested it myself and found that it often also does slightly worse than zstd -19 in compression ratio but faster (it was slower in one case on "uncompressible" input).

Compression performance vs "zstd -19" seems to depends a lot on actual input data in a very unpredictable way. I'd assume the benchmarks that they show are definitely somewhat cherry-picked.


having used zstd, it has terrible defaults, optimized for speed and low-memory. You need to change it's params (not just level and dict size) to get high performance.

probably somebody should use a coding agent to do auto-research to optimize params for each compression algo, while matching one fixed goal - time, memory or size


There is no way shipping alone explains the difference.

I would suspect the majority comes from higher (VAT) tax level; possibly also some indirect overhead because Europe is a more segmented market, and needs a bunch of localization at the very least.


US prices dont include taxes, EU prices do

£70 ($94) in the UK so is US prices exclude sales taxes $14 is VAT, what about the other $10?

GP says its $105 in the EU. I do not think VAT is high enough to explain that anywhere in the EU?


> GP says its $105 in the EU

It's not? Maybe there is some dynamic pricing going on, but it's 79 Euro (~91.75 USD). Which would be close to the US price with Hungary's VAT (27%, which is I think the highest). Then add stuff like copying levy onto it and you'll land on roughly the same price, even with a lower average VAT. https://en.wikipedia.org/wiki/Private_copying_levy

Maybe they compared the model with the base to the one without? The one with base is more expensive in the US too though.


There were leaks before the Manhattan project even concluded:

https://en.wikipedia.org/wiki/Klaus_Fuchs

Sure, non-US nuclear arms would have most certainly happpened either way, but exfiltration did presumably at least accelerate the timeline.


I'd argue the US only got dragged into lots of conflicts because of its role as de-facto global hegemon.

A lot of wars it did not really start; if you blame the US for Vietnam (instead of the french), you also have to count South Ossetia, Abkhazia, Transistria and Chechnya 1 & 2 on the Russian side.

The industrialized nations that actually did laudably little warmongering in the last half centuries are (somewhat ironically) the core WW2-axis powers: Italy, Austria, Germany, Japan; despite having plenty of "opportunities" to start something (e.g. over Königsberg/Silesia, Dalmatia, Kuril islands, ...)


> you blame the US for Vietnam (instead of the french)

I'm all for saying french caused a lot of death and destruction, but the Vietnam war was clearly created by the US during the 1954 Geneva Conference. Also, US-British involvement in France politics during the Indochina war (under Truman: Roosevelt didn't had to deal with Dixiecrats and give racist idiots a new enemy) probably made the 'sale guerre' last way longer that it had to.


My view is that France tried to hold on to Indochina unjustifiably for almost a decade after WW2.

So they get the blame because the US could have hardly escalated a proxy war without proxies.

Not saying America is blameless though, but I do blame France for setting the whole thing up basically (split Vietnam), and for all the wrong reasons, too.


You ought to remember that 40% of french voted communist in 48. Without heavy involvement of CIA and MI6 to curb the movements (including assassinations), the sentiment would probably have grow, which would have stopped the war (also, now that I think about it, it also means that the soviet are even biggest culprits on this as they did kill more french leftist in 5 years than everyone else combined). The more centrist government also would have tried to get out by 49 without the heavy British investment, and later USA support. And during the peace deal in 54, the milk toast also known as Mendes France accepted a first peace deal that the US then refused, and the division of Vietnam was probably US decision too.

"...laudably little warmongering.." - forgetting China here perhaps?

Not counting China (and South Korea) because they joined the "industrialized" club much later than the rest, but so far it is looking good, and I'm somewhat hopeful because their demographics make war an increasingly risky proposition.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: