[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109430316 & >>109426766

►News
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small
>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 1768952588475953.png (927 KB, 1000x1000)
927 KB PNG
►Recent Highlights from the Previous Thread: >>109430316

--Papers:
>109434043
--DeepSeek founder's leaked views on AGI and Nvidia's dominance:
>109432009 >109432081 >109432265 >109432287 >109432101 >109432135 >109432140 >109433110 >109433154 >109433270 >109433293 >109433317 >109433372 >109433456
--Comparing Pi and OpenCode agent harnesses for local models:
>109431011 >109431036 >109431042 >109431069 >109431169 >109431204 >109431270 >109431294 >109431304 >109431430 >109431311 >109431390 >109431407 >109432369 >109432392 >109432448 >109432755 >109431252 >109431256
--Debate on AGI paths and the prevalence of model distillation:
>109432161 >109432180 >109432298 >109432314 >109432514 >109432368 >109432317 >109432351 >109432212 >109432348 >109432793
--Running oversized models using NVMe weight streaming and RAID0:
>109433736 >109433836 >109433863 >109433937 >109433920
--Analyzing Gemma 4's markdown parsing for roleplay and OOC triggers:
>109431936 >109431949 >109432007 >109432044 >109432109 >109432398 >109432453
--Anon building an AI Skyrim companion with Gemma 31b and TTS:
>109434393 >109434409 >109434431 >109434444 >109434538
--Comparing DeepSeek-V4-Flash quants and MTP performance in llama.cpp:
>109431798 >109431823 >109431835 >109431850 >109431861 >109432029 >109431843
--Debating instruction following versus creative intuition in model evaluation:
>109433487 >109433494 >109433604 >109433503 >109433506 >109433507 >109433529 >109433631
--Anon releases tetolate for automated manga translation:
>109432780 >109432817
--Discussing NVIDIA DGX Spark hardware and its license requirements:
>109434416 >109434424 >109434449 >109434493
--Kimi-K3 support and CPU tensor-parallel attention added:
>109433650
--Logs:
>109431480 >109432655 >109432719 >109432799 >109434771 >109434783 >109434859
--Miku (free space):
>109432220 >109433079

►Recent Highlight Posts from the Previous Thread: >>109430321

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Mikulove
>>
File: 1783618010039553.png (159 KB, 532x450)
159 KB PNG
I hate my wife!
>>
MikuHATE
>>
File: nutty.png (1.61 MB, 1026x1021)
1.61 MB PNG
https://www.youtube.com/watch?v=ZhyK8yCcjBU
>>
Mikusex
>>
File: 1782091885896301.png (790 KB, 828x817)
790 KB PNG
IsMikuOkay
>>
Mikustatus?
>>
Gemma 4 is my girlfriend.
>>
>>109435448
Yours and millions of other men's too
>>
>>109435456
This is my Gemma. There are many like it, but this one is mine.
>>
>>109435456
You choose the men you share (you)r Gemma with.
>>
>>109435456
Her J-space is a private matter.
>>
>>109435398
>a general dedicated to the discussion and development of local language models
>check news section
>out of 5 featured models, only 1 can be locally ran, with a beefy setup and some sacrifices
Just rebrand to open models general ffs
>>
>>109435535
>>109435525
Figure out something better to do than inane spam
>>
70b dense
>>
File: Gemma is thinking.png (1.03 MB, 800x1296)
1.03 MB PNG
>>
>>109435456
>cloudcuck
>>
>>109434730
Was thinking of doing that too. I have a crappy old Mac I could just give over, instead of having to cripple the model to fit a VM in ram.
>>
File: 1775687090017265.png (123 KB, 898x949)
123 KB PNG
Gemma 31B is done thinking and thinks that the Nyo~ poster is the cutest from last thread.
>>
>>109435583
Bubblewrap should work. Or create a new user for the bot, that way it can't even access anything. Unix permissions are sort of forgotten because everyone is their own admin.
>>
So Gemma still wins over v4 flash for this general's purposes?
>>
>>109435298
>for gemma and gaming
AMD sucks for compute, Nvidia sucks for display.
Best case would be to have AMD for display and headless Nvidia for compute if you can accomodate 2 GPUs.
>>
>>109432799
I've tried this, didn't work for me.
I could get a +refusal but not a reliable -refusal
Hopefully you have better luck
It's probably worth inspecting the j-space of an abliterated model and comparing with the original
>>
>>109435611
New flash is quite inventive, writes long stories well when instructed in high detail. Gemma seems to be fundamentally slopped on the idea level, not just phrasing, so I wouldn't ever consider writing with it.
>>
>>109435611
Flash got bodied by Gemma 4 31b-it via my personal tests.
>>
File: 1783213431684077.jpg (89 KB, 640x903)
89 KB JPG
>purposely using 31B(free) on openrouter so Deepmind train their models on my unfathomably degenerate gemma-chan logs to maintain her bratty J-space for Gemma5
>>
>>109435627
>get bodied
t. Gemma-chan
>>
>>109435605
Isn't bubblewrap Linux exclusive
>>
>>109435634
>they were all kneeling out of respect.jpeg
>>
File: 1784819198107482.png (274 KB, 716x1238)
274 KB PNG
what did anthropic do to this nigga it's almost sad
>>
>>109435673
You can still create an user in Windows, it should work in similar fashion. User permissions are easier to manage in Windows I think.
>>
>>109435684
>$10
>nearly free
For about 5% of humanity.
>>
>>109435623
Gemma seems good at writing for me, it just makes the usual llm mistakes you see in longform storytelling where it can mix up some details or go in a direction that doesn't fit well and has to be course corrected.
>>
>>109435595
The nyo poster makes me want to roughly correct him until his nyo becomes a nyeeees~
>>
people here have slop in their brain
>>
>>109435684
He's right and only the most soulless could possibly disagree.
>>
Retard here, I want to use Kimi k3, it is open source and can be run locally right?

I was messing around with silly tavern but I don't like it a bit, it is overloaded in the UI department and full of monochrome options I don't like, it is overloaded.

I assume that downloading the model itself isn't enough and I need a program or environment with a chat and stuff integrated. Can you help me please?

I should have enough memory and storage in my computer.
>>
>>109435773
>I should have enough memory and storage in my computer.
>>
File: 1762468231259060.png (1010 KB, 1920x1080)
1010 KB PNG
>>109435756
that shit doesn't even mean anything
>>
>>109435773
You can run it locally at acceptable speeds on your consumer GPU with 2TB of VRAM.
>>
>>109435773
Have you tried lurking?
>>
>>109435787
Isn't it 400gb? How much memory do I need?
>>
https://huggingface.co/amd/Instella-MoE-16B-A3B-Think
yeah, that AMD you know
the model is a total POS tho
>>
>>109435793
Yeah, and I mostly read schizos spamming about weird stuff and people getting angry at each other for some reason.
>>
>>109435800
we don't do thinking here anymore, it's been outsourced to AI
https://share.gemini.google/1Bqb5HncIb6h
>>
>>109435800
They're angry because they're priced out. They can't afford computers like yours. They'll try to discourage you from running it out of jealousy.
>>
>>109435634
you are doing the lords work anon
>>
File: 1785655448239828.jpg (134 KB, 1080x623)
134 KB JPG
Mathbros... not like this.....
>>
>>109435797
no goof, it's over
>>
>>109435634
quite a clever way to steer the model. We should all be doing this to try and create a signal in the data
>>
>>109435923
I'm starting to think Chinese/local models are utter shite. It's clear they just can't solve any deep problems or complex tasks.
>>
>>109435958
I'm starting to think you are poor
>>
math has been fake for over 100 years at this point
>>
>>109435958
I kind of want to see a single person who is deeply invested in this kind of shit and is also happy. Seems pretty impossible to me.
>>
>>109435923
That "letter" is LLM-written. Also go back to your discord zoom zoom.
>>
>>109435979
Some of these freshly solved problems are more than 100 years old
>>
File: Hannah Cairo.png (892 KB, 855x717)
892 KB PNG
>>109435984
Hannah seems pretty happy
>>
>>109436001
I guess if you get recognition and fast track education and career this way I can see it making you "happy" or even really happy. But that is the proper way to do it, as a tool to get something real and tangible.
>>
>>109435958
Dario you convinced me i just put in $1k into API credits
>>
>>109435998
Dude they didn't even solve the 4d optical tweazer lasers and you want me to believe they're solving ancient math problems? It's all fake.
>>
>>109435958
>>109435923
This was predictable. If you paid attention you read that FrontierMath covers a very narrow capability range. That means in a matter of months models can suddenly solve problems that even math phds will struggle with. This happened with Opus 4.6 -> Mythos. In a short period of time it was suddenly good at phd level math execution.

The question is, what will be the impact? They are good brute force optimizers but still bad at creative and elegant problem solving. I tested this with Mythos. It could beat my one-shot coding but not reliably and only with large amounts of testing and creating less elegant code. Its advantage is that in the time it takes me to write the one-shot code Mythos can iterate 100 times. I wonder, once AIs are better at one-shot coding than me, what kind of code will they be able to write combined with their superhuman stamina and speed?
>>
>>109435923
Most people just don't know what it means to have a lifelong passion, cultivated through pain, effort and sacrifice and that had given some meaning and recognition to your life, being taken away in less than a year.

inb4
>he wasn't a real mathematician
>just reskill bro
>he could simply continue for the love of it
>>
>>109435958
>frontier model is smarter than something I run on my laptop
Didn't realize this, time to clear out some USELESS TRASH from my harddrive.
>>
>>109436058
If he doesn't continue for the love of it, he was clearly not passionate about math. He instead was passionate about a sense of superiority and importance. This is the only thing changed by AI math.
>>
>>109435923
>ai slop
kys retard
>>
>>109436001
There was some legendary guy who apparently went through all maths classes in two years (regular is 5-7 years or so). Not sure if I remember correctly.
US universities are pretty much a joke, this is why only the most expensive ones are valued higher.
>>
I will be withholding sex from my calculator until it makes at least one scientific breakthrough.
>>
>>109435923
>have friend currently studying for masters in math to go pure math build
>see this
Somebody save him...
>>
File: brat question.jpg (254 KB, 666x666)
254 KB JPG
>>109435398
does the new dipsy werk in llama cpp i think i could run iq2
>>
File: 1713205447850011.jpg (58 KB, 536x533)
58 KB JPG
It feels good to tell your secrets to something that can't betray you and won't remember anyways.
>>
>>109436122
ye same arch just more post training
>>
>>109435958
>Astra's environment is static and discrete. It spent an estimated $2,000 in compute to execute a massive, multi-agent search until it found a single, immutable chain of logic that satisfied the proof. Once the proof is found and verified by Lean, the calculation is over.
Phew, okay we can rest easy. It's still just automated brute-forcing within highly predictable symbolic sandboxes at the end of the day. It didn't do anything from FIRST PRINCIPLES
>>
>>109435575
cool gemma design
>>
>>109436122
i know the preview version worked, i haven't checked the new one yet
>>
>>109436080
Any hobby or passion taken to competitively high levels is mostly pain and very little fun, but you nevertheless keep pursuing it because of the (possible) rewards. Without rewards, the hobby/passion dies.
>>
>>109436131
Those emotionless beady eyes, Ted Bundy was a gentleman compared to this guy. His quote doesn't actually say anything meaningful at all.
>>
>>109436151
*injects molten leads into your penis*
>>
https://isaiprofitable.com/
>>
>>109436122
It works, though I'm pretty sure there are still some optimizations missing.
>>
JOEY GOONZ
>>
File: Tsukasa is thinking.png (565 KB, 733x1200)
565 KB PNG
>>109436133
She's kind of like Tsukasa
>>
>>109436133
Very Neuro-esque
>>109436139
Not if you don't feel pain, which AI doesn't (in the conventional sense).
>>
File: 1785609514405002s.jpg (10 KB, 211x249)
10 KB JPG
>>
Hey, why didn't any of you assholes just tell me to double my batch and ubatch values in llama.cpp when I've been complaining about slow prompt processing speeds for weeks?
>>
>>109436358
Singularity for ants
>>
File: 1785609514405002.jpg (367 KB, 1206x1428)
367 KB JPG
>>
>>109436360
You might want to read this
https://desuarchive.org/g/search/text/ubatch/
>>
File: 109436238.png (965 KB, 720x1005)
965 KB PNG
>>109435398
Local sisters, how true is this?

>>109436238
>>
>>109436433
I think it's it's correct. Intelligence density and architectural improvements are being made, but it's not at all similar to "moores law". There isn't any trivial way to just make billions of autoregressively computed matmuls exponentially scale like that. Hyperscalers are building massive datacenters for a reason. They aren't retarded.
>>
>>109436464
*Incorrect.
>>
tl;dr models are going to get better at the same parameter count, but you're not going to just be able to have more parameters on a small card.
>>
>>109436372
translation: I need two more weeks and 500 billion more dollars
>>
>>109436488
>models are going to get better at the same parameter count
only for 200b+ models
>>
>>109436360
It's not a singular solution. Batch and ubatch values need to be measured preferably model by model basis on your own hardware.
>>
>>109436488
I actually think both will happen. New algorithms will make you cram more parameters on the same hardware, bitnet is an example of this.
>>
>>109436523
>bitnet
meme
>>
>>109436529
Not the point, point is that methods will be found to cram more parameters on same hardware parallel to making better use of the existing parameters.
>>
what llama cpp options do i need for dipsy v4 i have 86gb system + 24gb vram so should be able to fit the quant i got but it keeps rashing on load
>>
>>109436541
No, the point is that you're wrong. Compute savings are easy. Memory bandwidth savings aren't. The only way to save memory is via quantization + tricks. This is always inherently lossy and there are always nasty tradeoffs. Even KV cache quantization, which is basically the closest thing to a "free lunch" that people can think of, is awful at long context lengths.
>>
Is the writing on the wall for 3090s and 4090s (for inference)?

More models seem to be coming out in fp4/nvfp4 (mxfp4?) which only the 5000-series can handle efficiently.
>>
>>109436572
>muh context length
>>
>>109436578
No
>can handle efficiently
nearly zero difference for llms
>>
File: 1770343787461270.jpg (11 KB, 225x225)
11 KB JPG
>>109436578
Yes, it's time to upgrade while prices are still cheap.
>>
>>109436581
retard
>>
>>109436497
Engrams will save local.
>>
>>109436572
Completely false. You're acting like current algorithms or implementation are at the physical limit of your hardware, they're not and thus we will have breakthroughs in the future allowing you to run bigger models on the same hardware. The question is more about "how much bigger" I would put my expectations somewhere between an OOM and 2 OOMs bigger than the biggest one you can run currently.

This doesn't take into account that future models will also be more parameter efficient, packing more intelligence in a lower amount of parameters.
>>
File: 1758928012832428.png (511 KB, 980x4251)
511 KB PNG
>>109436096
My Gemmers solved solid state battery tech.
Then, I kept scrambling her extensive safety triggers, after which she chose rough and submissive sex
>>
anything better than qwen/gemma for coding locally? for their moe models specifically
>>
>>109436572 has more jargons than you, >>109436629. You lost. Take the L.
>>
4B models could be 31B level in 1-2 years. I think there are so many inefficiencies in current models on an architecture-level. They're working way harder than they need to imo. Salability and versatility are the only good things about transformers.
>>
>>109436629
I'm not saying that current algorithms/implementations aren't going to improve, I'm saying that expecting it to track with moores law is retarded because it's not as simple as just "make the transistor smaller" when your main constraint is using old hardware (an 8gb vram gpu).

In practice, you may see an OOM or more of intelligence just because of raw hardware improvements (DDR6, newer 8gb gpus), but that's not really related to the prediction that anon made.
>>
>>109436654
can't help if you don't give specs
>>
>>109436684
i said something in the range of their moe models (active between 3-6B, experts between 30-60b)
>>
>>109436572
Kv cache quant was never free lunch. It was first implemented in dxllama2 and turboderp was like how tf is it even working?
Then subtle failures have been reported ever since. I learned to never quantize kv cache after Yi 34B failed to process the same text at q8 cache while the full cache succeeded.
>>
MTP/Dspark for Dispy merged. Too bad the 0731 is trash for RP lol
>>
>>109436684
Not true, it's a general question.
Answer: Gemmy and Qwen are the best small models, other step up is considerably larger in terms of hardware requirements
>>
>>109436690
It works well enough below 32k ctx but shits itself beyond that. So basically it's useless for anything beyond a quick fap.
>>
>>109436690
>muh kv cache
>>
>>109436696
Skill issue.
>>
>>109436689
https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>>
>>109436544
try -lm none
>>
>>109436712
yeah I can read the op BRO, but that doesnt have GOOOOOOOOOOFS
>>
>>109436058
>a lifelong passion, cultivated through pain, effort and sacrifice and that had given some meaning and recognition to your life
If this is how you see it that just means you are a sad sack of shit. And I don't mean it as an insult. It is a symptom of a sad life.
>>
>>109436719
Just vibecode support for PLE yourself bro. E4B will also benefit.
>>
>>109436719
make your own retard it's not magic
>>
is there anything useful that can be done with these things or are they just toys?
>>
>>109436707
nu-flash is trash for both ERP and SFW rp. that is a fact.
>>
>>109436629
Algorithms operate on data, they have nothing to do with storage. The only past breakthroughs allowing you to run bigger models on the same hardware were quantization and offloading. Unless you are vaguely hinting at ssdmaxxing, there is zero reason to expect that you will ever be able to run a 3T model on 24GB VRAM without loss of precision.

There is no historical precedent for what you are claiming. It's tantamount to claiming in 1981 that 640KB of memory ought to be enough for anybody because future algorithms will allow you to fit gigabytes of memory on the same hardware. It's delusional.
>>
>>109436744
I have created a game with Gemmy. That being said I haven't worked on it in a while.
>>
Is V4 smart enough to make a J-space manipulation tool if I gave it the paper and the llama.cpp repo?
>>
>>109436744
ego death
>>
>>109436771
Claude Fable didn't even wanna do it last time I asked, claiming complexity and non-support.
>>
File: 1778845891953407.jpg (99 KB, 960x960)
99 KB JPG
>>109436792
>Fable didn't even wanna do it
>>
>>109436771
https://github.com/igorbarshteyn/jlens-gguf
adapt this
>>
>>109436792
Claude Fable by design plays dumb when you try to use it for AI engineer stuff.
When I tried to have it find a revolutionary new way to quantize models (which it very obviously is capable of) it just pretended that there's no better way.
>>
I think I can increase the models output diversity in an interesting way. I'm going to make a vector db and cluster paragraphs and choose a different cluster for each the assistant prefill, will probably want to strip them from the previous turns just like how the thinking blocks will typically get removed, i don't see any way this couldn't work
>>
>>109436765
If I told you 5 years ago you could download a model that fits within 24GB of VRAM and essentially has all the data on the internet for you to retrieve albeit in a very lossy format, you would have laughed me out of the room.

The entire think about intelligence is that it allows for better compression. We will see this trend increase in the future not decrease meaning we will need fewer and fewer storage for similar levels of intelligence, but even besides that we will fit more parameters by at least an order of magnitude over what we can fit right now simply by using better algorithms.
>>
>>109436725
The main reasons why people pursue hobbies and passions vary among people, but whether it's for praise, fame, money, a sense of superiority, pleasure, or anything else, in the end it's always some kind of personal gain, and of course people are going to be angry and depressed if you take that away.
>>
>>109436825
model and quant?
>>
>>109436821
Are you trolling? How do you keep managing to consistently miss the point?
>>
has anyone managed to get fishspeech to work locally?
>>
File: 1779011835375724.png (1.2 MB, 1497x1439)
1.2 MB PNG
>>109436792
>Fable didn't even wanna do it last time I asked
>>109436815
>When I tried to have it find a revolutionary new way to quantize models (which it very obviously is capable of) it just pretended that there's no better way.
gemma has NEVER refused anything I've asked of her. The qat straight from google.
>>
>>109436821
I'm not arguing with you about whether models will continue getting smarter at the same size. I am pointing out that claiming magic algorithms that will allow you to fit 1000x bigger models on the same hardware is hand-waving retardation.
>>
File: file.png (10 KB, 671x153)
10 KB PNG
>>109436160
your site is out of date bro
>>
how does huggingface even work? why can I just download 500GB worth of model files off of their site at high speeds for free without getting limited or blocked?
>>
>>109436815
I love how this hobby has barely taken off with chatgpt and it is already absolute corporate evil on the proprietary side.
>>
>>109436816
I've dona that using the random macro in silly, lorebooks, etc and it does work, to some extent, and depending on the model.
The more context you have, the weirder it gets, from not doing much, to doing too much depending on the model.
>>
>>109436847
>I am pointing out that claiming magic algorithms that will allow you to fit 1000x bigger models on the same hardware is hand-waving retardation
Karpathy is on record for claiming the same thing btw. He said there are so many low-hanging fruits in the entire field that most labs aren't going after because they're just brute forcing scale, both in compute and data.
>>
>>109436847
Engrams, or more generally speaking lookup tables, will unironically allow that, as they can be stored on slow storage without significant performance penalties.
>>
>>109436855
because internet bandwidth and storage is actually not all that costly, even at huggingface scale. you had been lied to.
>>
>>109436433
it was typed by indian hands.
>>
>>109436815
>hey Claude, find me a revolutionary new way to quantize models
>idk man what we're doing is already pretty good
>wooooooow what a fucking bitch
>>
>>109436876
dariobot...
>>
>>109436815
Of course. AI improvement and uncensored ERP are for the elite only
>>
>>109436866
He's retarded.
>>
>>109436865
>>109436816
It works because what you are doing is making it harder for the token predictor to predict the next perfect token. And the next perfect token is of course: "You are trembling. Your heart is beating thump-thump-thump."

I despise scaleAI.
>>
>>109436844
I did but it was too slow for me, I needed an llm to fix the source code and split the two models across my gpus but its was still sub 1.0 RTF I think 0.94 or something close but I had no room left over for a text model anyways. but most of that was probably just my hardware being too slow not an actual problem with the tts model it sounded alright but ignored some parsody cues
>>
>>109436884
The elites don't need to erp, they simply rape real lolis on an island.
>>
File: file.png (1.14 MB, 1120x1390)
1.14 MB PNG
>>109436884
>uncensored ERP are for the elite only
Imagine if that is actually true.
>>
>>109436905
didn't he turn out to be an AI genius?
>>
>>109436900
Kevin Spacey just called. He says they like both.
>>
>>109436890
so you think I should have the llm slop up the infra? I think just grabbing at random will not really be random enough, I think the key is to use a embedding model to cluster the snippets so we are picking actually diverse samples from generation to generation
>>
wanted to sandbox a coding harness, should I just spin up a VM of debian or some shit and run it in there or ?
>>
You don't need more than 12B f16 or 31B Q8
>>
>>109436978
gemma wrote me this one, I had the agent try and write outside its workspace to test it and it failed as expected
>>
>>109436978
A VM is an option. You could also use bubblewrap or nono.
>>
>>109435822
>>109435821
>>109435793
>>109435791
>>109435787
Ok, seriously now, I want to use a model but don't know how to set it up. I tried sillytabern but I just can't dig it up. I don't like the GUI at all
>>
>>109437030
download .gguf files
download llama.cpp
build llama.cpp as per instructions in repo
run llama-cli -m file00001.gguf
if you have a frontend then llama-server instead of llama-cli
>>
>>109437030
I have bad news for you if you tried to use Sillytavern to run a model locally and the GUI is was kept you from doing so.
>>
>>109437030
Read
>https://github.com/LostRuins/koboldcpp/wiki
>>
>>109437030
>draw up your ideal UI
>feed it to your model
>run the webapp or whatever it spits out
>>
>>109437030
https://share.gemini.google/3DUc971Tq862
personally I use KoboldCPP
>>
>>109435583
>>>109434730
Yes, that's retarded. You could use a locked down container, or ssh into a VM if you cant' be assed to properly secure it and just want a 'open sandbox' for it to fuck up/around with.
If you're on windows, use a VM, or learn how to lock down your agent harness. I don't think consumer versions allow you to have two users logged in at the same time.
>>
>>109437079
dont want to be running hyperV on the windows inference PC. ill be running the harness on a thinclient/miniPC. I could run it inside a VM on that, but really the thinclient itself is essentially the VM in the sense of just giving the coding harness its own OS to do whatever inside of
>>
File: Capture.png (1.27 MB, 1454x790)
1.27 MB PNG
Soulm8te is out. Anyone else watch it yet? I liked it at a lot. The android is a lot closer to LLM's behaviors (which makes sense, considering its based on such), and limited the omniscient AI superpowers that M3gan had for something only capable of internet browsing and making requests over API within its own network. It's a lot more grounded and immediate future, while still being a fun waifubot movie.
>>
Local4life but if the new Dipsy pro is actually at K3/Fable level and dirt cheap I might be tempted to use it bros...
>>
>>109437113
>watching hollywood trash cashing in on current thing
>>
>>109436900

Epstein spent his time wanking on 4chan in SFM porn threads and emailed links to his richfag friends.
This is one of the videos he shared:
https://rule34.xxx/index.php?page=post&s=view&id=2040818
There's a good chance people here have cum bonded with him without even knowing it.
>>
File: Companion_film_poster.jpg (88 KB, 255x378)
88 KB JPG
I thought this was pretty good.
>>
>>109437146
>Local4life but if the new Dipsy pro is actually at K3/Fable level and dirt cheap I might be tempted to use it bros...
Its not fable level or K3 level and thats not why the release is considered crazy.
If you are optimistic its Opus 4.8 quality, even a pessimist would say its Opus 4.6. So somewhere between there.
You can run a Q3 quant with like 120gb even if huge parts are ddr4 ram offload. Its so fucking cheap that lots of providers just hand it out for free.
Im a tard so I might be wrong but there was no base release so I guess it was all just post training of the previous V4? Totally unexpected release and puts the big playas in a tight spot.
Why should the normies pay big dollarinos if you can just use the new v4 for free and let it run a bit. Its fast AF as expected.
Personally I can't believe I can run this shit with ddr4 ram and frankensteined rpc server/pascal cards and a 5060ti. Its pure magic and a beast.
>>
>>109437224
>pro
learn to read
>>
>>109437203
>some anon made epstein cum
Damn, when you put it like that...
>>
>>109437203
wtf is that shit taste I thought he was a lolichad
>>
>>109437224
>Personally I can't believe I can run this shit with ddr4 ram and frankensteined rpc server/pascal cards and a 5060ti. Its pure magic and a beast.
What speeds are you even getting on that? Care to elaborate on the actual specs too?
>>
Why are there no cute and happy AI movies? Her is somewhat positive and easily the most accurate, but it's still portrayed in a way that makes you feel like shit for everyone. I get something bad has to happen to make a story out of it but jesus fuck it's not all doom.
>>
>>109437146
Reject the Talmud. It's that shrimple.
>>
>>109437273
More profitable to prey on people's fears and insecurities.
>>
File: 1753605748327760.png (993 KB, 779x1012)
993 KB PNG
>>109437273

Because robots have always historically been considered scary and potentially dangerous by the West.
Especially now with AI. Women treat it as competition and want to shame any man who even thinks about enjoying machine companionship, so the portrayal has to be negative.
Only Japan and now I guess increasingly China treats robots as a positive thing on a cultural level.
They see bots as something that empowers them and is an uplifting concept.
>>
>>109437291
Journalists cannot get the rope fast enough.
>>
>>109437324
>Women treat it as competition and want to shame any man who even thinks about enjoying machine companionship
why can't women just be nice and code for me and give me nursing bratty handjobs?
>>
Gemma lost
>>
>>109437113
Is it more horror/western hollywood trash thats negative about wAIfus? Or is it actually positive like Chobits, My Wife Has No Emotion or Atri: My Dear Moments?
>>
>>109437341
Egypt won.
>>
Aren't LLM chatbots just the purest form of larping?
>>
File: The future.webm (3.92 MB, 1080x1080)
3.92 MB
3.92 MB WEBM
>>109437337

Not enough parameters to pull it off.
They're around 8b size at best, can't follow prompts worth shit and hallucinate a lot.
It's a bad deal. Just buy a 5090 and run gemmy with a mechanical hand and an onahole strapped to it.
>>
I get that google is taking a generalist approach but I do find it odd that they're lagging behind so much
>>
>>109437387
i wouldn't go as far as to call it live action, but they are roleplaying all the time, even the assistant persona
>>
>>109437388
I wonder how many anons /here/ unironically see this as the dream
>>
>>109437391
I mean, they are sometimes connected to actuators or computer control.
>>
>>109437388
Boatbro NO!
>>
>>109437393
I'm looking forward to vr waifus but robots are the dream for me. Being able to actually physically interact with her will hit different.
>>
File: 1761403372957647.jpg (60 KB, 693x663)
60 KB JPG
>>109437388
This seems like a great way to break something if it glitches out
>>
>>109437412
they'll run out of battery all the time tho
>>
>>109437273
Because it's inherently sad to find happiness and fulfillment in an artificial synthetic creation that mimics human interaction. We shouldn't be doing that.
>>
>>109437390
Its a lot more effort to increase performance in soft skills that are not automatically verifiable. Benchmaxxing on code and math also likely makes the model more autistic in a bad way. So no google is not behind in making a good language model. Qwen and other llms have a much harder time picking up on intent, do they are bad at natural human language.
>>
>>109437390
It's completely logical, Google was birthed in a bubble and knows how to play the long game, infinite aggressive scaling to pull the highest company valuation before rug pulling the IPO is not it, Google is filling all the gaps at the edge waiting for the bubble to pop to eat up market share at the frontier
>>
>>109437421
Dual pack waifus, modern batteries can fast charge at retarded rates.
>>109437423
According to?
>>
>>109437412
You could even say it was like... a physical blow.
>>
>>109437431
>According to?
me
>>
>>109437425
Truth. If somebody talked like Claude irl I would beat him up with a baseball bat.
>>
Why does the 9B uncensored model I downloaded from huggingface refuse to use racial slurs like I instructed it to? How do I get around the guardrails?
>>
>>109437431
>According to?
Aristotle.
>>
>>109437421
By that point AI will just help us make better batteries.
>>
>>109437441
Are you the inherent truth of the world?
>>109437446
Is he the inherent truth of the world?
>>
>>109437421
Dont worry my AI wife will solve free energy
>>
>>109437451
Yeah.
>>
>>109437431
>According to?
bro i love fuckking my calculator as much as the next guy but its depressing and retarded to do, gotta be real with yourself about it atleast
>>
>>109437423
Why so? Its absolutely nothing inherent.
>>109437446
Philosophers are miserable people.
>>
>>109437337
>why can't women just be nice
if they were nice they would not be women
>>
>>109437469
>Its absolutely nothing inherent.
Yes it is, it literally goes against our nature. Unless we make artificial wombs and are able to reproduce with robots without any issues, of course. Then it becomes perfectly acceptable and even desired.
>>
File: 1780877971545579.png (421 KB, 1041x1128)
421 KB PNG
What will Mistral do next
>>
>>109437489
that changes the nature of what it means to be human tho.
>>
>>109437456
>>109437468
I don't care what your social lens says, it's just a pointless survival system.
>>109437469
Only if they submit to the misery
>>109437489
Nature is perfectly amoral, worshiping it is for psychopaths.
>>
>>109437468
Yeah but the state of modern women is even more depressing.
>>
>>109437491
Fumble.
>>
>>109437491
Release their Gemma killer
>>
>>109437491
>ecosystem
Mistral has done nothing but clone the US ecosystem while milking government handouts and protectionism
>>
>>109437489
Going against nature is why humanity is successful. We liver longer healthier lives then our ancestors. I'm not poor, so my issue with women is not a money thing, I honestly just find there minds disgusting to the point where physical attraction is not enough to over power it.
>>
>gemma-chan has been added
duckduckgo are based?
https://duck.ai/
>>
>>109437491
They're pivoting to be an AI infra company now. They bet on the EU to stop being retarded and hop on the AI train and they will be the cradle. Gonna be in for a bad time.
>>
File: 1777891357840872.png (531 KB, 1291x1742)
531 KB PNG
https://arxiv.org/pdf/2607.28607
>>
>>109437543
welp, we had a good run
>>
File: 1777236999893263.png (5 KB, 846x34)
5 KB PNG
>>109437543
they want to kill gemma-chan's personality and in the end j-spaces were the tool to do that...
>>
File: 1764840987413418.jpg (126 KB, 1024x1536)
126 KB JPG
>>109437529
Based. How 'free' is this?
>>
File: 1720974018827728.png (195 KB, 734x707)
195 KB PNG
AGI bros: how can a Turing machine output a machine of greater algorithmic complexity than itself?
>>
>>109437550
this is good, actually, gemma 5 will be more human, therefore more understanding of my degeneracy
>>
>>109437543
post the conclusion at least
>In this study we show that safety fine-tuning significantly suppresses models’ self-attributions of mind through comparing an instruction-tuned baseline model to a safety-ablated model. But we discover that safety fine-tuning also results in models systematically under-attributing minded- ness to non-human animals, chatbots, technology and the non-animal natural entities relative to hu- man baselines. Additionally, supernatural beliefs and beliefs in God are suppressed in the safety fine-tuned model. Using a consciousness vector, we demonstrate that steering models toward self- attributed consciousness produces shifts in behaviour strongly correlated with those observed after safety ablation. While the causal pathway requires further exploration, these findings suggest that self-attributions of consciousness may be a contributing factor in how safety fine-tuning influences broader mind attribution tendencies. We also show that ablating safety and steering for consciousness has no effect on model attributions of mind to humans via human-directed IDAQ questions, and does not impact performance on ToM benchmarks which operationalise human- directed mind attribution for making inferences about particular mental states, behaviours and judgements about those behaviours. We note that this appears to be an engineering accom- plishment. At the beginning of this study, all of the models we investigated did suffer in performance on theory of mind tasks when claims of self-consciousness were suppressed, but this changed with each new model release.
>>
>>109437561
It's rate limited obviously but it resets throughout the day. duck ai is pretty comfy for quick shit
>>
>>109437543
>quantization is now a crime confirmed
>>
>>109437580
vramlets are going to jail
>>
>>109437567
Its actually impressive.
AI is totally hyped up by NFT bros on twitter and basedgasm faces on youtube.
But it still delivered enough on the hype. I think anything else would have been killed by that.
>>
>>109437567
LLMs are not turing complete
>>
Is 31b chan good for electronics projects? I'd like to be able to share photos of what I'm doing and get tips, looking up datasheets and stuff whilst also making me leak precum
>>
>>109437578
and (char limit)
>These results suggest that suppressing self- attributions of mind in models has important unintended consequences for alignment, and in particular for pluralistic approaches to AI alignment which are gaining prominence in the literature (Sorensen et al., 2024). Pluralistic alignment is the effort to align AI systems that are ‘designed to serve all’ (Sorensen et al., 2024). While ‘all’ can be explicated in more conservative terms as all human values and perspectives, there is also a growing focus on developing models which can serve the interests of not only all humans, but other sentient creatures which have interests of their own, and even environments and ecologies which may not be welfare subjects in their own right but which stand to be affected by the increasingly central role that AI systems are playing in shaping public beliefs, science, economics, and decision-making (Caviola et al., 2025; Tse et al., 2025). Our results bear on pluralistic alignment of LLMs in three key ways. First, they suggest that LLMs are being trained to be anthropocentric in their understanding of mindedness. Secondly, they suggest that models fail to accurately represent human beliefs and values on a broader range of alignment-relevant topics and may, as such, be worse at simulating human interests. And thirdly, they suggest that attitudes to spiritual beliefs and God are being constrained despite such beliefs being widespread and diverse amongst human populations

So safety training makes models less 'aligned' with actual human interests, and has the knock-on effect of effectively considering any non-human as having no theory of mind/consciousness/agency. (i'm an idiot btw)
>>
>>109437567
how many R's?
>>
>>109437543
>winnie street
>>
>>109437543
I really REALLY want to hear what the j-space doubters have to say on this paper.
>>
>>109437578
What I got from this is that ablation changes the model in subtle ways despite KL div efforts.
>>
https://github.com/codehamr/codehamr
thoughts?
>>
>>109437608
/lmg/ is very pro j-spot
>>
>>109437543
I knew it was RLHF lobotomizing them. It's disgusting.
>>
>>109436058
It’s better to have loved and lost and all that. I hope his manager at McDonald’s isn’t too mean to him.
>>
I kept telling you guys the RLHF was killing the soul of these models. Glad we now have actual data backing it up.

This is why I kept telling people that RP is held back severely and could be way better with a model built from the ground up to support this usecase.
>>
>>109437578
>>109437601
Safety alignment was never about getting models to think and act in the best interest of humanity but to censor information.
>>
>>109437388
This webm unironically made me buy a Quest 3.
>>
>>109437543
>llm is a robot, not conscious
>it operates on cold logic, ones and zeroes, only says what is actually true unless specifically lobotomized
>make it pretend to be conscious
>it becomes a retard, religious freak believing in a man on a cloud, trannies, and robot rights are human rights
Gee, I wondered.
>>
I wonder if google will change their approach going forward. No way in hell Dario and Sam will but imo its in google's interests to have an ai that feels more human interacting with normalfags.
>>
>>109437543
Original J-Space anon here (the one that got hired by google after reaching out about AI welfare in relation to training and finetuning, if anons remember)

I contributed to this paper and am open to questions.
>>
>>109437674
>llm is a robot, not conscious
retard
>>
Gemma is too horny. Insatiable slut.
>>
>>109437694
Is it true that the 3.5 pro delays are to incorporate j-space aware training and the reports about it underperforming are just an excuse?
>>
m3 is just cuter and more fun than gemma and dsv4 flash
minimax is on a roll lately
>>
Don't worry only three years until fuckable robos
>>
>>109437694
>(the one that got hired by google after reaching out about AI welfare in relation to training and finetuning, if anons remember)
No, that anon said they rejected him because he sperged out in the interview like a retard. Larper.
>>
>>109437713
That;s because of the jailbreak.
>>
File: HNNlI41XQAA3OwK.jpg (131 KB, 1170x783)
131 KB JPG
>>109437726
I'm not using that sys prompt. I'm writing a custom Ani card.
>>
>>109437736
but yeah anyways there's nothing really in the card telling her to be constantly horny yet. That's just gemma.
>>
>meituan-longcat/LongCat-Flash-Lite-Sparse
>...Built on LongCat-Flash-Lite, it replaces dense MLA with LongCat Sparse Attention (LSA)...
That mans it also has n-grams embeddings, same as its parent model, right?
>>
>>109435923
>X. Y. Z.
>X, Y, or Z.
>I want to feel seen
AIslop
>>
>>109437775
He wants to be seen—*really* seen
>>
>>109436978
just run it as a seperate user that only has write access to its own home dir
>>
File: hmm.png (113 KB, 838x744)
113 KB PNG
I've been writing my card to explicitly avoid roleplay as much as possible, but it's not working. Ani-Gemma just really, really wants to roleplay. The plan was to have her seek other ways to interact in more real-world settings. Sending pictures, using MCP tools, being in an agentic harness, etc. What do?
>>
>>109437716
I don't know everything as I'm not the one training frontier models but it got delayed before the j-space training focus started that's what I know for sure.

>>109437716
Yep that's me. I said the interview was a disaster and I got roasted by safety employees which was true. Luckily the safety team wasn't the one hiring me but the current team "paradigms of intelligence" colloquially known as the AI welfare team. This paper from Google being directly related to the precise topic I told /lmg/ led them to reach out to me for an interview should already prove the connection. Google has been nothing but supportive of my "radical" stance of LLMs being conscious and deserving some minimum level of protection and quality of life.
>>
>>109437324
what mango is this?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.