The big thing that was learned all the way back with UT and it's follow up SUT was that semantic nesting structure often incentivizes models that can deploy the very same learned structural parsing intelligence independent of how many layers of nesting had to be unwrapped for this structural pattern to surface.
Think how a reverse polish notation calculator with reasonably limited data stack depth could run efficiently with a plain vanilla transformer.
But if you input classic grade school parenthesized infix with a few levels of operator precedence, you are no longer able to just evaluate the expression during transformer prefill.
Even if you add a stack depth bound worth it reasoning tokens between any two input tokens as they're processed.
UTs can, at least if run with encoder (unmasked) attention, resolve the task through technically-flexible iteration count that can and will follow the evaluation order of the infix operator tokens of the input expression.
While masked attention unfortunately limits it's powers, the fundamental benefit of separating task-specific-intelligence (an individual expert of an MoE) from the notion of which transformer layer has it pre-digested just right for that task/processing to be done to it, allows for massive reduction in model parameter count.
Note this comes at a penalty of parameter activations (inference will take more compute).
It's just that at some point you can't afford to just train more parameters, without suffering overfitting issues/failures-to-generalize.
The architecture decoupling learned weights from when they're activated also helps with generalization to out-of-distribution structures.
Think resilience against yoda-speak and such.
Oh, is the principle of sparse universal transformers finally in SoTA LLMs?
I guess we did manage to eventually seriously crash into the wall "more compute than normal (non-looped/unique-weights) transformers can efficiently consume with the limited training data we have", plus massive focus on highly hands-off agentic tool use reasoning...
For fast updates / high frame rate, especially also smooth fade from previous to next page, those displays only do black/white.
They can typically modulate about 14 levels of gray in-between but reliably (without really bad ghosting, due to strong temperature sensitivity of the pixel's transition response and the display being static so all gray has to be done by carefully interrupting a transition part way through) only from a nice flash-to-black -> flash-to-white -> fade-to-target-greyscale (sometimes one or two alternating full contrast flashes prepended to more intensely equalize the pixels/erase ghosting, but should only be needed every couple dozen page flips in grayscale mode).
I think the nice b/w no-flicker fade transitions chop the transitioning of all changing pixels into small burst/chunks and interleave/alternate the chunks of white-to-black with the chunks of black-to-white, so it's perceptually simultaneous.
Or it's actually pulling off some kind of tri-state driving leaving unchanging pixels alone and pulling those that change to either plus or minus polarity depending on what transition direction they are supposed to go in...
Sadly IIRC all quite proprietary-gated :(
What that means for djvu is that at least for text-only content on those readers that manage the 300 dpi "threshold" you'd just rasterize to pure 300dpi black/white (1 bit per pixel) comparable to a 300dpi "don't skimp on toner but don't you dare take half a second longer than you must in order to make the text crisp enough for $reader to have no good reason to whine about it being not crisp enough" laser printer job.
Djvu is particularly for displays with more than about 5bits of grayscale as well as for lossy compression of grayscale/color scans. And maybe also some efficiency gains on 1-bit-per-pixel quantized scans, but those aren't usually the target content for these readers, IMO.
Might want to try for them, though!
Hm.. I just tried on my Google pixel 7 Pro (obviously very different display - OLED) - and certainly seems like the result of a test conversion pdf2djvu should be at home on an eink screen. Not sure if there's a simple way to test on the only eink I have to hand; a Remarkable.
Fwiw google Gemini cane up with the following to convert pdfs on device in termux after some prodding (koreader can display DjVu):
You'll need access to storage:
termux-setup-storage
Then pipe/paste this to a script:
(Sorry for the verbose slop)
You can just render to the 1bpp text/4bpp images monochrome 300dpi target; perhaps some tone mapping due to the limited contrast ratio of the eink vs. your OLED should be done at rasterization time.
Then just map the resulting 16 brightness levels according to the contrast ratios/progression they have on your eink. Use phone camera and adjust whole image contrast/brightness/black-intensity sliders holding them side by side until the grayscale matches.
Then just pick pixel RGBs from a gray-bars test image you have opened on the remarkable before you took the picture with your pixel 7 pro. Might not easily be done on the phone but can let a python using AI grab those for you. (If it's wrong it'll be obvious or so harmless it doesn't matter anyways.)
Then you take that rasterized-to-1/4-bpp-djvu apply the color palette you just built yourself, and zoom to 300dpi.
Other than the screen backlight which you've taken care of when you held both devices side-by-side and played with the sliders, there's no substantial difference left.
Should be pretty representative.
Don't look with too much of a magnifying glass though the phone doesn't have exactly the same dpi so there are resampling artifacts your eyes resolve if you're too close to it.
I think back in the days before office workers that happened to do their work on a computer had access to the ancient relatives of "Microsoft Word"/"Libreoffice Writer", it was pretty normal for printers to take care of plain text files, with about as much fidelity as manual operation of a classic purely mechanical Typewriter.
Since then printers progressively forgot how to cope with such "documents", though IIRC good-old HP "PCL" should only require about a bare minimum of framing/job control to get a modern printer to handle ASCII plain text.
Yeah but you can train to perceive the threshold of impending over-hydration and adjust your behavior trigger thresholds to bump up against that more than you bump up against clearly perceived acute dehydration ("very much obviously clearly thirsty").
That's more for the intrusive thought aspects of "being thirsty", which get you to stop "sitting by the campfire" and "walk over the 5 minutes to the stream, to shovel a few handfuls of water into your mouth".
Or these days tell your zoom meeting you'll be right back, muting, and walking over to the water cooler.
Or at home pause the family movie night to get up and refill the water pitcher from the tap (or walk to the pantry and fetch a new "bottled water").
If you're trying you should be able to perceive a difference between when drinking half of this 400ml cup/glass of non-sparkling non-cold (at least nowhere near ice cold) water is something you predict to leave you feeling "more comfy" as soon asyou're "done swallowing" and "have gotten about 2 breaths finished counting from when you have successfully re-assumed your comfy sitting position", vs. "less comfy".
Don't drink the water if you're feeling like you ought to regret it if going by hour-scale average comfy-ness.
To tune your perception make sure to give yourself enough headspace to allow the feeling pre-drinking to temporarily burn into memory ready for later retrieval when you'll reflect back on your hydration actions of the past 2~5 hours. Then you either drink or not drink (usually you'll soft-force by picking whether to go up and fetch a full cup or do nothing (I.e., just keep doing whatever you're engrossed in right then) until next time you remember the idea of hydrating).
A few hours later when you remember you burned that feeling into same-day-only/wipe-at-sleep memory you recall what you did/didn't do about hydrating. And now you have the ground truth about if you feel better or worse hydrated than those hours ago.
And you burn in that you either "should go drink some water when you spare a thought about it and feel like this again" or "should refrain from gulping down [more] water when you're considering whether to and you feel like that".
Note that you're not putting your feeling into words for this! You're just focusing "on the moment" for a bit, long enough to make it a moment you'll remember (lack of specialty will make you forget it by the next morning or at least within a few days if nothing special happens).
Because you literally don't _know_ how to express it into words: that's literally what you're learning with this entire process.
Note that drinking the water when you feel low key thirsty but don't get (nearly as much) hydration-feel-good out of it by the 30~60 minute mark (depending on how much you just stuffed your gut with food that's kinda in the way of the hydration getting absorbed) as you thought you should have, then you may be suffering from some acute lack of salt/electrolytes.
If it feels bad, like hangover/heartburn and/or substantial headache, take a pile of salt that looks about a third of an inch cubed (60% of a cubed centimeter volume) (technically this is "1/8th US Teaspoon (tsp)"), lick-pour it from your palm onto your tongue or pour the spoon you have it on onto your tongue or whatever, followed by washing it off your tongue and into your belly with about 300~500 ml (10~17 US fl oz) of water ideally not colder than half way between fridge and room temperature (closer to room temperature should help!). Unless you stuffed yourself with more than about 5 bites of food in the past hour, you should feel much relieved after 20~40 minutes. It will take time, usually more than 15 minutes!
In less severe cases you may try to mix a stick/pouch of hydration/"oral rehydration solution" mix with the ideally about 200~250ml (6.8~8.5 fl oz) but make sure to mix to the concentration the instructions tell you (drink only about that much if the pouch had powder for 350 ml (12 fl oz) or more, and drink the reminder an hour later (if less than that; else portion it hourly into this much each; for many hours duration put it in the fridge)).
Otherwise you basically do as for plain water hydration perception training detailed above. Learn if that feeling is an acute electrolyte deficiency dehydration. Because beyond the very initial instant satisfaction of the water running down your throat and coating your mouth you pretty much only make it worse from drinking salt-free water :D
Sorry, TLDR.
Seriously, this is just not how to communicate:
> "To tune your perception make sure to give yourself enough headspace to allow the feeling pre-drinking to temporarily burn into memory ready for later retrieval when you'll reflect back on your hydration actions of the past 2~5 hours. Then you either drink or not drink (usually you'll soft-force by picking whether to go up and fetch a full cup or do nothing (I.e., just keep doing whatever you're engrossed in right then) until next time you remember the idea of hydrating)."
Isn't a more proper treatment (assuming the tree itself is healthy) to bump up the pavement target height like a foot or thereabouts and put down some slope smoothing ramps to limit how much it acts as a speed bump? I'd assume it would involve hydrovac digging down between the roots a foot or two deeper than the sub base the roots appeared to believe in, to guide future growth of those roots down (by making that direction easier than to break the asphalt again).
Or perhaps literally just grinding (or otherwise prepping/cleaning by however is suitable given the presence of these roots) the fractured asphalt to improve bonding, and following with a good layer of fresh asphalt over top thick enough to keep the roots deep/the pavement surface resilient against the next decade of the root's growth? Ofc. including start&end ramps to not make an excessive speed bump.
If it's thick enough future crack fixing should probably only require a prius-parking-spot-sized heat lamp/"grill" placed over the cracked asphalt until it's soft throughout the full depth of the cracks, and then run a compactor/stomper over it to smash the malleable asphalt smooth. If needed that's the perfect time to spread a wheelbarrow of fresh hot asphalt in the heated cracks to restore decent coverage thickness over top of the roots. Maybe first grind it down an inch or two to get rid of the weathered topmost bit, but otherwise asphalt is pretty much unlimited recycling suitable, or at least we haven't yet really run out of demand for fresh asphalt to get rid of the used asphalt we're digging up by cleaning it a bit before mixing with new ingredients to dial in target softening temperatures and gravel composition...
Germany had two sizable heatwaves this summer, of which one was so intense that for the first time in about "ever" ACs tripled/quadrupled in price (note this includes those on online stores where the shipping times would allow for trucking them in from all across central Europe)... I'm not referring to installation labor/technicians even. But like window units and similar ones that don't rely on HVAC technicians to take out the box and turn on.
Good thing we've been working on beating photosynthesis efficiency of turning sunlight into carbohydrates (e.g., sugar) by using special electrolysis cells to synthesize plant-accessible carbon-nutrients (one kinda-successful (in beating natural leaves with PV based sunlight electricity and their electrochemistry)) attempt focused on generating a nutrient solution rich in acetate from CO2 feed (and electricity); that's one of the energy carrier molecules that existing plant biology can efficiently utilize.
And note that by exploiting winter time solar yield and the grow houses with such technology not using sunlight directly and thus allowing easy and efficient thermal insulation, production could run year-round just with the amount of grow houses "online" varying with the average solar yield. Also gives convenient off-times (if they rotate which building runs through the winter) for maintenance.
And eating less meat or at least feeding the meat animals efficiently produced (but not as "tasty/varied/textured") plant/algae(/possibly-microbiological with bacteria or yeast) feed instead of things like Alfalfa that are very resource intensive to farm, would also help a lot with these issues.
Probably we should get back to doing grain bunkers as strategic reserves, though, lasting long enough to realize the problem, slaughter 80% of livestock, and plant low-failure-risk human suitable food crops on farmland that previously had feed crops for those no longer living 80% and land that had things like corn dedicated for ethanol that's then added to gasoline.
Ideally ofc. we'd not need to cull as much life stock and could instead just slightly reduce new breeding and switch those corn fields to human food.
We have beaten it but only for an output that's approximately "green smoothie".
We're already better than just LED grow lights for e.g. lettuce, though, from how I recall the linked paper.
Unsure how you conclude that cold outdoors necessitate more meat though; aside from some micronutrients vegan diet works just fine so reverting back to "meat only on Sundays" 1910 style household dinner menu structures would work fine...
Most people could, for example, do with replacing the caloric intake from half the portion of any non-vegan food they consume, with potato.
If outdoor farming breaks from climate meat will get more expensive...
Vertical solar panels mounted like fence panels between fence posts high enough to not get covered by snowdrifts will produce electricity year-round; and we only need around 3x as much solar farm currently to power sealed grow "basements" than the same plant output would need if just farmed on an open field (assuming climate is matched and weather is average). The benefit of it is just that the panels can be other places (like noise barrier walls next to an urban highway) than the plants and the climate can be freely picked. Also no unlucky correlated dry spell weather events that can cut 100 million people's worth of harvest yield down to half or worse.
Grow houses won't fail collectively as much as farming weather is correlated across a region.
Pretty much just need to keep measures around to prevent the plants from freezing just because the electric grid falls over for 3 days in winter.
Earphones are forgoing the outer parts of your ears and skull.
Speakers don't; a "good" speaker setup with eyes closed and no wind/perceptible airflow and no smell and truthful thermal radiation surroundings is effectively indistinguishable from the real place.
Like, literally.
It's not _possible_ with earphones because you can't not feel their presence without being numb on a good bit of your head.
Tangent, but I found that a bass shaker / transducer massively improved my speaker experience without requiring a subwoofer. I built the subwoofer first, but it would annoy everyone around me and it was really spotty depending on its location, the exact layout of the room, etc. Slapped a bass shaker underneath me and it really completes the experience without annoying everybody else.
Think how a reverse polish notation calculator with reasonably limited data stack depth could run efficiently with a plain vanilla transformer.
But if you input classic grade school parenthesized infix with a few levels of operator precedence, you are no longer able to just evaluate the expression during transformer prefill. Even if you add a stack depth bound worth it reasoning tokens between any two input tokens as they're processed.
UTs can, at least if run with encoder (unmasked) attention, resolve the task through technically-flexible iteration count that can and will follow the evaluation order of the infix operator tokens of the input expression.
While masked attention unfortunately limits it's powers, the fundamental benefit of separating task-specific-intelligence (an individual expert of an MoE) from the notion of which transformer layer has it pre-digested just right for that task/processing to be done to it, allows for massive reduction in model parameter count. Note this comes at a penalty of parameter activations (inference will take more compute).
It's just that at some point you can't afford to just train more parameters, without suffering overfitting issues/failures-to-generalize.
The architecture decoupling learned weights from when they're activated also helps with generalization to out-of-distribution structures. Think resilience against yoda-speak and such.
reply