It was previously thought impossible to create a 1bit model that didn't become lobotomized to the point it can barely string to words together but now thanks to big daddy google its possible to make 8b models fit in 1 gb and 27b models fit in 4 gb and benching them has shown they perform 90-95% as good as fp16 versions of their models.
>>109333826Benching them as a bloody benchod basterd beach?
Wild fucking claim. I heard good things about Ternary Bonsai, but that's too much.
>>109333883How are you uncs so out of touch?
>>109333826>27b models fit in 4 gbdoubt.jpeg
>>109333826>its possible to make 8b models fit in 1 gb and 27b models fit in 4 gbno it isnt
>>109333826>-1, 0, 1>1 bitNot sure if nu g is genuinely retarded or if it’s bait
>>109333826But what for? The model training companies will just increase the amount of parameters and you still won't be able to run shit
>>109333990>>109334006Have you guys been living under a god damn rock for the past week?
If your balls are in three boxes, can you pull out in time before she realises its both your wifes son and you DPing her box?
>>109333826>1 bit>look inside>three possible states
>>109334295I specifically mentioned one bit AND TERNARY models you fucking retard
>>109334315Yeah but whoever made the image you posted, presumably some faggot jewtuber making a thumbnail, is giving me a fucking headache with their retardation.
>>109334329it's one ternary bit, duh
>>109334355"Bit" is a portmanteau of "binary" and "digit". It can't be a trinary binary digit you fucking turkey.
>>109334045>Setup: little-coder harness via the harbor adapter, all 89 tasks of terminal-bench 2.0, single attempt (k=1), 40-turn cap, temp 0.2. RTX 5070 Laptop 8GB, i9-14900HX, 32GB RAM, CUDA 13.1. Runtime is PrismML's llama.cpp fork (stock llama.cpp can't load the 2-bit kernels).>Results: Ternary-Bonsai-27B at 2-bit scored 7.9%, Qwen3.5-9B gets 9.2% and Qwen3.6-35B-A3B gets 24.3%, both as per-trial means from their k=5 runs. The 1-bit Bonsai never produced a number.Usecase for retarded 27b models?
>>109334441soa tit?
soviets were right https://en.wikipedia.org/wiki/Setun>>109334441its called a trit
>>109334763this
>>109334622>look we made 27b perform worse than 9b!Great result, I can see why people are hyped
>>109334763for the impatiens anons, ternary is closer to the euler number 2.71... that it makes it optimal for many mathematical operations
>>109334622>doesn't mention the retard who posted that on reddit ran the bonsai model with 16k context instead of easily using 40k context >wonders why it spilled its spaghetti when it forgot whatever the funk it was doing half way through the prompt Why are you people like this?
Question: how does compute scale with <8bit models?Do they get expanded to bytes to perform computations or do they use bitwise operations?I get that RAM is scarce but it seems wasteful to throw away 87.5% of your compute power if that's indeed how this works.
>>109333826>1 bit>{-1, 0, 1}eggcelent
>>109334802it's equal to euler's number
>>109334860no ternary is 3... which is obviously closer to e than 2- just more dificult to use
>>109334878yeah, ternary is 3so is eas is pi
The paradigm is fundamentally flawed until you replace or stop depending on back-propagation.