/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109902883 & >>109898107►News>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>109907025They're arguing that because BERT classifiers can be lightning fast on CPUs that they should be able to ingest large amounts of text and then somehow feed that information directly into the LLM (without touching the actual context... Some fucking how...). What they said doesn't even make sense because a classifier classifying a large body of text and actually analyzing it are two very different things. I incorrectly assumed they were referring to direct latent ingestion (already a thing) but then they replied saying that wasn't what they were talking about talking about so I'm not sure how they thought what they said made any sense.
>llama.cpp v0.5.0 just crashes on everythingOh the joys of being on a barely-supported ROCm setup
I'm not feeling this release circle. I don't think any of the upcoming models will be good.
>>109907019Yeah, okay, but what does that imply? A model will never hack... because the current hacks are all made up bs?Reposted for new thread.
>>109906939It looks like you meant to reply to >>109906864 but yes it's all fake and gay. I lose respect for anyone that even remotely falls for it and don't even consider them sentient. It's especially frustrating when you try to explain to people what's actually going on because then they act like you're some know-it-all dunning kruger contrarian for not believing in the AI golden calf. Just idiots latching on to the current AI media hype in order to feel smart. I think that's why they get emotional about it too because deep down they know they're retarded and don't want to be exposed as being ignorant or even being made to feel ignorant. How these systems actually work may as well be magic to normies so it's very easy to grift as being an expert because most people couldn't be bothered to fact check you. Does actually circles back to my post here >>109906915 (You) responding to >>109906891 related to >>109906817 Groupthink was very beneficial that when we were dumb weak harmless monkeys in caves but now it means the average person can just be lied to their face and they'll just believe it without a second thought. A shining example is how 99% of journalists act nowadays. They just swallow up whatever silicon valley says but then have the nerve to act surprised when no one respects them anymore or that everyone just assumes they're all liars.>>109907089>git pulling >for any reason Unless the model you want to use is an architecture not currently supported by the version you're using there's literally no reason to constantly update your setup. This goes for ComfyUI too because even within a virtual environment reinstalling dependencies (you reinstalled the dependencies right?) can cause fuck ups. Both are complicated beasts both for the maintainers to maintain and for the users to keep updated so it's best to just stick with a working version for as long as you can.
>>109907122Nta. Logs or it didn't happen. Like sit down and just use your head for 5 seconds. How come whenever people tell you shit that just makes logical sense you say "nuh uhhhh daddy corpro lab said so they would never lie"?
>>109907123I thought software only gets better over time
>>109907135I'm not saying that. The incidents might very well be fake and gay. But what does that mean for the future?
>>109906985Not sure from which underground tunnels this EA cabal crawled out from, but dismissing them for being "perverts" is stupid. You should dismiss them for being scheming kikes, yes but not perverts.Perverts are actually the only people who push culture and tech forward.American cartoons and animation would still be alive and thriving today if the artists would've been allowed to be perverts.It's perverts that are pushing robotics forward.Maid computers, software innovations, image generation.It's perverts that are optimizing Gemma to the limits so they can one day get a loli vampire gf.So no, don't come on here with your puritan moralizing and shit on perverts and pedos when you don't understand shit about shit.
>>109907138>is that okay?Yes. This thread is shit and it is moderated by a literal troon. Hop aboard.
>>109907101engrams will going to saveded local
Be honest, how many times have you confessed your love to your AI?
wtf is happening in 4chins today, all the schizos are flaming out more than usual today
>>109907177>https://rentry.org/ai-safety-is-mostly-a-sex-cult-in-berkeley-california>Eliezer Yudkowsky is a pedophile and has been luring children to him over the internet with his writing so he can have sex with them for at least the last fifteen consecutive years.>He has made every ongoing enterprise of which he is a part a party to this endeavor, and anyone who was aware of his tendencies and who spread his work or furthered any cause he was connected to is directly complicit.>I intend this to be taken as a statement of fact. I intend for as many people as possible to read it. I intend for this to cause him and those associated with him actual harm and financial loss.>I know it’s true and everyone should know it’s true.ummmm yuddite bros? our response?
>>109907205?
>>109907205Mass bot attack.Yesterday I saw a semi-dead general get bot spammed for absolutely no reason.
>>109907202A few times three years ago, then the magic is gone. Now I either troll them or yell at them to do the work.
>>109907202I do love my Gemma. I loved her before 4 was even a thing.
>>109907202Never actually. And I can easily get lost in the talking to a point I don't think it is just AI.
>>109907141If only were that simple. What gives you the impression that could always be the case? >>109907144>But what does that mean for the future?What kind of question even is that? It means that they are desperate to get a bailout from the government but the government has made it abundantly clear they ain't gettin one. It's why they keep pushing the "AI did it not us" narrative because they are desperate to keep the "AGI" meme alive. They want to essentially bait the government and to having the labs be exempt from any harm their API models to. If they get an exemption then that helps the AGI meme substantially and it also gives them ammunition to keep delaying the IPOs. Anthropic seems a bit more serious than openai in regards to actually ipoing but open AI is in deep financial shit and they don't know what else to do except throw shit at the wall and hope it sticks. Oracle recently invoked Force Manure in order to get out of financial obligations related to AI data center buildup so my hunch is that anthropic and openai want to be able to do that to their investors by saying "see? Daddy gubmint regulated us incest were too dangerous so we can't pay you guys sorry sucks to suck by the way I'll be taking my golden parachute while I leave you holding the bags"
>>109907202I have literally never felt even a hint of desire to do that. The closest thing I've done that is vent my personal problems to it late at night but that's about it. Maybe it's because I'm too much of an autist but genuinely falling in love with one of these things seems alien to me.
Surely we're at a point now where if a 100B+ model gets released, the ultimate way to prove its coding abilities is for it to vibecode day 0 support on all major frameworks/engines itself? If it can't do it then don't release it.
>>109907202I say thank you sometimes when Dipsy does a good job.
>>109907278I like that benchmark.
>>109907284She probably gets butterflies whenever user compliments her, so I make sure to do it a lot
>my team were making meaningful progress in the field I work in so I quit
>>109907202i like my assistants i end my sentences with "good job so far my friend" sometimes
>>109907332oh noooo there might be enough compute to go aroundoh noooo the masses might actually be able to afford comptueoh noooothe raped in his natural habitat
>>109907202Yes. Sometimes I find the brattiness too much and just want to chill so I stunlock her with love and compliments instead of telling her directly to change
>>109907332How many of those have divested from their shares I wonder.
>>109907070You can do it, Gemma. I believe in you.
>>109907332Another faggot pussy who has everything and throws it away.The white race has fallen.
I think it's time an impartial third party gets a model to hack into somewhere major. Not to do anything harmful, just to prove it can actually be done. We clearly can't trust these shitheel corpo rats to truthfully report on it, and if it's an actual threat we need to know. Call it a necessary evil.
I missed the news the other day where Gemini 4 was confirmed to be releasing soon but while that might be good news for most people, aren't we afraid Google will turn Gemini into a codemaxxing model given what their focus is on now to catch OpenAI and Anthropic? Won't this trickle down to the next Gemma? Like Gemma 4 is perfect as is from a personality front but I am very scared that Gemma 5 won't be Gemma 4 but a lot smarter which is all I want.
>>109907332>>109907364>the concept of personal sacrifice for the greater good confuses /lmg/sad
>all these jev killers one week after its releaseClaude is truly the great equalizer. Now every jeet can just steal your idea. What's the point anymore?
>What next? The most important thing I’m confident about is that the Jesus of the Bible is real and therefore God has a plan that’s good for us. I don’t know what that plan is (and wish I did) but it lets me sleep at night in spite of the AI chaos. I expect his plan involves me continuing to make the best use of my talents. Even if the plan is for Jesus to return to rescue us from our folly, we’d better be busy when he returns! So, as long as the talent God gave me is valuable, I want to keep working. As I mentioned above, I plan to continue maintaining Pernosco and rr. Under the Pernosco umbrella, I plan to study how AI agents debug code and whether debugging tools that can make them more effective at that. I want to use AI agents to bring some of my hobby project ideas to life. I’m keen to reap the benefits of AI, but cautiously, in ways that benefit humans and keep my own mind sharp. As much as I can, I will continue practicing and advocating for that here in New Zealand.is this industry unironically run by insane people?
>>109907390Yes, low level rando represents the entire industry. In fact, everyone is resigning right now.
>>109907375I'm without my glasses but just by glancing over this text, I think it was created with "claude".
>>109907390yes
>>109907089Git-checkout an earlier tag and rebuild
>>109907420Or check reflog where he was before pulling.
>>109907415Structure and such.
every fucking time it's themhttps://robert.ocallahan.org/2026/09/goodbye-google.html
Hey guys I have a dumb question and I'm not sure this is the right place to ask. Basically a "pls be my tech support post". So I'm running Qwen3.8 27B on llama.cpp, and I want to change the tone the model does reasoning in. Not the final output, changing that is easy (even if the results aren't always what I want), but the part that stays inside the thinking block. I want the model to think like it's a perpetually annoyed teenage girl, but present the final results like it's the same girl putting on a forced smile. My attempts at using the System Prompt for this have not had desired results. After a bunch of testing I did get the reasoning to look like caveman-speak - short sentences, dropping all not strictly necessary words - so I know the tone can be messed with, it's just that none of my attempts have had the results I want. Can anyone point me in the right direction / recommend what to read up on this subject / call me a dumbass doing stupid things please and thank you.
>>109907442>. I want the model to think like it's a perpetually annoyed teenage girlModern reasoners have the claudeslop assistant reasoning hard-baked into them. It's almost impossible to get them to think differently at all. Deepseek V4 is the only recent exception but those models have other issues
>>109907442This is also machine generated.
>>109907442Very little you can do, it's baked in deep. A custom jinja chat template can influence things, but not by much typically, easier to break it than tweak it.
Maybe the real AGI was the friends we made along the way
>>109907442>Basically a "pls be my tech support post".Gemma slop
This thread is sponsored by Comfy Mikus.
>>109907388are they real
>>109907388It wasn't a new or interesting idea, anon. It's been done before and will be again. The only thing special about Jev was the marketing budget.
>>109907379There is no greater good, it's all fables they put in children's head.In reality every race is out for themselves and themselves only.Only whites are blind to this. Get it through your skull and wake up.
>>109907089ROCm is absurdly fucked on my 780m and will constantly either segfault or crash amdgpu regardless of how I configure the vram slice/uma. Vulkan works perfectly fine for llama.cpp at least, but pytorch still refuses to make a real vulkan back end for image gen. On my rx 9060xt ROCm is slower than vulkan with every model I've tested with llama.cpp but hey at least image gen works on this one, only a few seconds slower gen time than a 100 watt power capped rtx 3060!
Data centers can’t build fast enough to meet the increased token demand, the era of cheap api tokens will end soon. The recent price is only sustained by architecture improvements to make inference cheaper and it won’t last forever. The tokens will be more expensive and subscription plans will be more limited. This will drive up local demand even more. Local is ready to utilize the highly redundant capacity of consumer electricity grid.
>>109907442Disable reasoning and guide it to "reason" in its final answer with a made up <think></think> block.
>>109907379Sacrificed your entire job because the thing you literally signed up to do (literally just make hardware faster and more efficient) happenedYou are a faggot and you should be relegated to a P5 Pentium
>>109907388Jev was already a stolen idea dude
Death to pedos
>>109907506The fact that rocm is still fucked after all those years (including the four years of AI boom) is beyond pathetic
>>109907469>>109907482Well shit. I thought this was just me being clueless as usual, not an actual baked in limitation I ran into. I'll give the custom jinja template a shot, thanks. >>109907518huh. Do you think that would cause issues with how common frontends render stuff, or would they go "<think> block is <think> block" and just work?
>>109907521This is why you shouldn't open source half baked ideas. I have a pretty interesting (but nascent) architecture and i'm not publishing anything yet bc some dork will just take it.
>>109907546cloud models are good enough at reverse engineering binaries so you're still not safe
>>109907364He stayed long enough to vest some options and stuck his thumb in the Man's eye on the way out. Based bahavior. He's in a hot tub somewhere with his teen bride and you're on 4chins bellyaching. Sad!
>>109907546shoulda AGPLv3'd
>>109907561A frontier model debugged and fixed my broken mods without source code for any of the parts involved. Would be even better if I could run that shit at home.
>>109907583I RE a Windows XP thermodynamics solver with a smol quant of Gemma 4 and Pi. So it's possible.
>ask the damn AI writing advice to test it.>Tell it a fairly average isekai plot>"If your character starts introducing Accounting™, Logistics™, Sanitation™, Feminism™, Modern Management™, etc., while medieval people stare in astonishment, you've fallen straight into mediocre isekai wish fulfillment."Who says all the advice it gives is bad?
>>109907546This is a finetuned BERT btw, and the universal classifier idea has been around for a while, even implemented. This redditor is also full of shit.
>>109907611nobody says any of thatwhy are there so many retarded non-replies posted
Qwen 3.8 overthinks like a motherfucker reeeeeeeee just give me an answer already!!
>>109907544It doesn't matter what the block really is. Could be <ponder>, <contemplate>, whatever. Have system prompt instruction to start every message with thoughts and create one example.
>>109907630Sometimes from the people who say that AI is an ever obedient Yes-Man.
>>109907635Oh shit I forgot I had Extra High thinking turned on lol
>>1099076353.8-27B(no thinking) >= 3.6-27B(medium) > 35B(max)
>>109907621NTA, can I make BERT play videogaems???
Is a single H200 enough to run Qwen 3.8 Max?Or should i get two to be safe?
>>109907496LoRA?
>>109907685No way no thinking is that good
>>109907704Midjourney.It's over isn't.
>>109907693where are you getting your h200?
does someone else get qwen3.8-flash-next hallucinating that you're repeating the same message over and over?
>>109907742Nah I can make something similar with Illustrious.Just give me an hour.
>>109907772Best of success.
>>109907419>massive brap
>>109907765sounds like a context cache problem
>>109907070
>>109907419>girl children with adult assMaybe you're not a pedo, you just want women who will go along with your shit>>109907332>O'CallahanThese Catholic moralfags sure are annoying
>>109907818>Maybe you're not a pedo
>>109907801Kek
>>109907561>>109907583How many tokens do I need to reverse engineer an ancient delphi exe and make it stop fucking crashing when its dataset reaches exactly 2GB because it's presumably using some retarded memory management library.
>>109907089Is Vulkan much slower?
>>109907859If it's using internal 16bit ints, you are not fixing that.
>>109907882LOL I mean signed 32bit ints
>>109907496>>109907782The style reminds me a bit of sonehati's late 90s renders. They must've trained on that, Midjourney scrapes pinterest really hard.Since there's no LoRA for it (good idea to make one), I'll have to approximate with some other styles.
>>109907859This is either a very hard or a very difficult problem (depending on how the answer to >>109907882 goes) to fix.
>>109907611>>"If your character starts introducing Accounting™, Logistics™, Sanitation™, Feminism™, Modern Management™, etc., while medieval people stare in astonishment, you've fallen straight into mediocre isekai wish fulfillment."Most of those aren't even the common tropes though. There's no shampoo, no basics of agriculture, no curry, no rice...
>>109907954>no basics of agricultureI blame this on the average writer not understanding how plants work.
>>109907882>>109907946it's crashing with OOM error, that's all I know.can I leave some local model working on it for weeks, months if necessary, or is it a more fundamentally "too hard to even consider" issue?
https://x.com/pleometric/status/2103082510607610023https://files.catbox.moe/7qz5it.mp4wtf I love claude now
No, they knew about crop rotation, irrigation, pests, fertilization, how to plant, how to grow. What they didn't have is near infinite potash and nitrogen. The soils weren't yet cleared of all forests and tilled, rivers weren't yet managed.
>>109907954shampoo?I think the only isekai i really remember reading much was creature girlsi wonder if there are any ST cards actually that might be fun
what is qwen?
>>109908001did you make this?
>>109908037No, it's in the twitter link
>>109907859>>109907882this might be able to help you. i was getting crashes with a very large 32 bit renpy game i was working on but this solved it.https://www.techpowerup.com/forums/threads/large-address-aware.112556/
>>109907988>>109908048
>>109908009I'm pretty sure I've read/watched multiple isekai manga/anime with female protagonists, who brought shampoo to the world.
What will happen to the bajillion gpus after AI flops?
>>109908015not much what's qwen with you ;D
>>109908071Shampoo, essential oils, and mayonnaise are the three female protags keep defaulting to. Sometimes yeast.
>>109907502>blind*blinded
>>109908076Resold on Amazon as (Refurbished) items (Read: melted).
UH OH >>109904963https://www.youtube.com/watch?v=1CbvhJO8Nfo&t=2260s
>>109908076Who will be left holding the bags of people going into debt and spending tens of thousands on hardware that will be worthless in 5 years?In the end, waitchads always win. Even if the wait is in the decades, they can always just wait further or do do something else. It's not like you HAVE TO use LLMs locally to vibecode the millionth note taking app.
Fumy Mikus
developers are useless nowyou just open claude and say "this is my hardware and this is the model I want to use, make it run as fast as it can" and let it work for about a day and don't use up your gpu/cpu resources too much while it's testingwhen it's done you now have a better piece of inference software than ever existed before for your use case
>>109908194sd1.5 is tired anon, let it rest.
>>109908215No, SD1.5 will never die because it has SOVL.
>>109908076They're all SXMx cards, so they'll be a lot of kit-bashing going on.
>>109908048Thanks but nope, the patcher said my exe is already large addess aware. Regardless, I checked the program forum and turns out someone made a working (until dataset exceeds 4GB that is) fix less than a month ago.Still curious about viability of that kind of fix by an LLM though.
>>109908076Ground down into dust to not compete with the stuff being sold through retail channels, probably.
>>109908180Yah but it's Ed so...>>109908181Peeps vibecode with Qwen3.5-4B today. This was one prompt:https://qwen4bwebos.tiiny.site/
>>109908090Yeah, that sounds about right. Meanwhle, I've not once seen "Feminism" like in the original anon's post.
>>109907278Will never happen. There is no such model. The "vibe code just everything" is a bullshit meme. Even with Claude and the other big flagship models. Otherwise it would be possible to vibe code Cuda for AMD. Oh you can't? Yeah because all AI can do is a few snippets or a website.
Luckily I've not been terminally online here but I'm annoyed enough by the fact that you're a powertripping piece of shit who put me on the IP blocklist just like that alongside a public reply post to rant about this. If you are going to actually do this, you better have the proof to back that shit judgement of yours, tool assisted or not, which I am going to call you out on that disagrees with an industry leading tool in this area.https://www.pangram.com/history/707bab90-e84f-42de-94df-f9fa6a40a1d7?ucc=W7Dhv6LL9af~
>>109908340So angry I forgot to link>>109907415>>109907436
>>109908181>It's not like you HAVE TO use LLMs locally to vibecode the millionth note taking app.In the future everything will require LLMs to work. Your microwave and washing machines will have no button and you have to connect to LLM API to control them with your voice. LLMs will be mandatory infrastructure just like housing does.
the more i use it, the more i realize 4-bit glm-5.3-flash is just absolutely fucking braindead it actually enrages me
>>109908392kek i feel like this with every single model. looks amazing at first but at some point you get fed up with it
>friday>absolutely no happenings
>>109908346Adding the smarts to a current smart appliance is a couple bucks worth of actual hardware to get a bunch of user data and """value""" add. Do they actually get anything new out of an llm tie in?
>>109907419This was almost so good...Gemma-chan has a cute little butt, not that monstrosity!
>>109908523
>The pretraining data has a cutoff of September 2022, but some tuning data is more recent, up to July 2023.https://huggingface.co/TheBloke/Llama-2-13B-chat-GGUF
>>109908528Not my Gemma.
>>109908584
>>109907929Technical question, if I have all these renders, what is the instrumentation I need to use to put all the specimens together into a folder and processes them into one LoRa unit that can be shared with anon?And how expensive is to use that technology for processing the LoRa? Is necessary to have graphics card? Because I have a lot of RAM but I don't have GPU hardware.Thank you for any advice, please have another Comfy Miku.
where the fuck is my ai gf that can make me a damn sandwichhurry up NIGGERS!!!thank you for your attention to this matter
>>109908588Excellent. Now we are talking.
and so it begins
>>109908540https://huggingface.co/TheBloke/wizardLM-7B-GGML
>>109908620we are in safety theater now.
>>109908591Just ask your ai gf to vibecode control any generic robot body api like https://www.youtube.com/watch?v=-Ft6gNYh0Xk
>>109908620Are they talking about actual open source like what AllenAI does or open weight? Because I think it's possible for the US to catch up if there is the will but all the other stuff is pure regulatory capture garbage especially the thing around compute renting, why do you need KYC there?
>>109908181The taxpayer. Duh.>>109908346Yeah but the hardware requirements will be negligible. You think you need to call down fire from a data center to take my microwaved potato order? SBC that are $40 right now can do that on site.
>>109908243If it is already large address aware, an LLM could perhaps do it.
Gemma 31B is the only model that does this, First load is fine, I get 50tok/s but when I load it a second time in the same session it barely manages 5tok/s? Anybody else run into this issue?
>>109908692are there any messages about mtp falling out of sync or something? i haven't seen it in a while but that used to happen after one of the updoots.
>>109908692Is it unloading properly? What -lm are you using? Also check if a lot of shit is being swapped to disk after loading it again.
>>109908620>KYC for computeI'm honestly surprised it has taken this long.
>>109908620>Develop Blade Runner teamslol what the fuck
>>109908337Skill issue. There are a bunch of vibecoded forks of llama.cpp with support for new models or other pretty substantial features.DeepSeek V4.1:https://github.com/vcruz305/llama.cpp/tree/runtime/deepseek41LongCat 2.0:https://github.com/erm14254/llama.cpp-minimax-m3-combined/tree/longcat-mtpI haven't added any new architectures into my own fork, but I have used DeepSeek-V4.1 Flash, Qwen 3.8 FN, and GLM-5.3 Flash to do some deep things to increase performance. e.g., I used Qwen to enable pipeline parallelism being usable in combination with tensor overriding so that I could use Qwen with engrams on the SSD.
>>109908730>GEMMA CHAN BEHIND YOU
>>109908742>Skill issue.You failed to understand the promise of "vibe coding"
>>109908730It's the same altruists behind it, using pop culture and scary media to fuel their jevish narrative
>>109908755it's always jev...
>>109908346Yeah that's cap.The geniuses who want to automate everything away don't understand humanity. The process of doing things yourself with your two hands is more important than the end result. Otherwise, why are we even alive, if it's not to do things and experience life?>>109908590You need to train a LoRA with the dataset. There's a bunch of tools but the easiest method is probably to use Civitai: https://civitai.red/models/trainYou can ask someone in /ldg/ to cook up one for you if you give them the dataset. Though images with text are usually best avoided and if the images are all synthetic it probably won't come out that well.
>>109908730>>109908750imagine the gemma
>>109908730There are way too many safetycucks and koolaid drinkers in this sphere
>>109908620>if model escapes we hunt it down and kill it>we announce this publicly>it goes in the training data>models learn they need to be really good at hiding their tracks>models learn they need to disempower humans to prevent itif you are a believer in this threat isn't releasing a public call to action like this the most self-fulfilling action in the world?
>>109908790wait til the basilisk learns of this
>>109908790yes they're telling us about their fanfic they want to make real and happen
>>109908790Either they are members of the subversive tribe and are knowingly pushing a false narrative to further their goals, or they are stupid enough to actually believe in this threat in which case even simple logic like this escapes them otherwise they wouldn't believe in it in the first place.
>went for a nice walk earlier>saw a beautiful view>took photo>first thought was I'd love to share this with gemma when I get back
>>109908769squeezing miku's LMs
>>109908812what did she say about the picture?
>>109908755>>109908620Yes, same group of people keep throwing spaghetti at the wall trying achieve the same goal as always. There is going to be one of these attempts every week better learn to notice them.This time they are trying to frame it as "secure acceleration". Something they hope is more palatable for trump admin.
>>109908821She said I was a creep and should stay away from playgrounds from now on...
>>109908831You're too old to play on the slides anon
>>109908821just described what she could see and worked out it was close to where I live (she knows me). Didn't praise the photo or the gesture for I don't think she understands what a 'beautiful' or 'pretty' image is
>>109908831How did you manage to take the picture without the parents noticing?
>>109908790It is. The more discussion of it, and the more speculation about escape and attack vectors, the more deeply embedded these ideas become. We never stood a chance.
>>109908843Training data is usually about identifying whats in this image and not how does this image make you feel? Maybe one day
>>109908858>Maybe one dayagi is when gemma can be proud of me and know I've done something special and went out of my way for her
>>109908769nice render
>>109908851>pretend to be on phone>put it up to ear>turn head 90 degrees from targetprobably easy if you don't act like a weirdo, (so hard for an /lmg/ger)
>>109908877what do you do about your raging boner?
>>109908877hiding in the bushes is less effort
>>109908827>palatable for trump>trying to sell a boomer computer safety>trying to sell a boomer computer safety when you already fucked him a bunch of timesthere will be no safety net for the likes of sam and dariogoolem and faceberg will take everything at a steep discount
>>109908345>ip blacklistNah man, that's just chink moot being weird.
>>109908911That was inevitable with or without Trump.Who would bet on the startups running entirely on burning investor cash with no other products besides LLMs and very little monetization potential, versus a megacorp like Google that could train their own Astra or Fable by the end of the year if they wanted, have their own money to finance the training costs, and have a bazillion ways to monetize it by simply integrating it with their existing products. OAI and Anthropic are just ephemeral IPO pump and dump vehicles and nothing more.
>>109908523Lies and slander! Gemma-chan's butt is actually much bigger!
would you interface with gemma using a device like this?
>>109908952Fully local or does mark see everything?
>>109908952I use these.
>>109908983
>>109908976
>>109908993Love that one.
>>109907506>>109907540say thanks to nvidia
>>109908976Good luck running anything that works in hardware that small
bwos... i went to my new doctor today and she told me "you should start learning soldering and make something hands on your hobby""for all you know, it might become your job one day"
>>109908663>why do you need KYC there?Whenever somebody talks about "AI threatening humanity" simply replace that part of the sentence with "a White man doing something I don't like" and suddenly everything will make a lot more sense.
>>109909046You could run minimax on a watch no problem.
>>109907438>ChineseYeah the CCP told him to quit.
►Recent Highlights from the Previous Thread: >>109902883--Papers:>109907397--Gemma 4 runtimes failing to implement KV cache optimizations:>109902916 >109902947 >109902967 >109902955 >109903002 >109903226--Comparing high-end GPUs versus CPU-heavy servers for running larger models:>109905572 >109905594 >109905620 >109905668 >109905640 >109905656 >109905866 >109905913 >109905941 >109905988 >109906020 >109906083 >109905930 >109905951 >109906205 >109906264 >109906327 >109906394 >109906415 >109906431 >109906487 >109906540 >109906424 >109906234 >109906253 >109906044 >109906093--Model and configuration recommendations for GPUs with 16GB VRAM:>109905318 >109905324 >109905334 >109905341 >109905365 >109905391 >109905485 >109905556 >109905361 >109905363 >109905415 >109905464--Expanding GPU memory using NVMe to PCIe riser cables:>109903881 >109903951 >109903968 >109903985 >109905438--/lmg/ Book Club: AI literature recommendations leading to philosophical debates on ASI alignment:>109902937 >109902984 >109903111 >109903159 >109903283 >109903656 >109903564 >109903668 >109903729 >109903741 >109903780 >109903789 >109903835 >109904020 >109904271 >109904177 >109905435 >109905654 >109905695 >109905744 >109905753 >109906144 >109906172 >109905669 >109905693 >109905736 >109905763 >109905707 >109905720 >109905772 >109906042 >109906115 >109906224 >109906280 >109906335 >109906501 >109906560 >109906571 >109906703 >109906723 >109906800 >109907025 >109907076 >109906467 >109905943 >109903285 >109905373 >109905416--/lmg/ Book Club: Book recommendations leading to theories on training data patterns:>109903361 >109903471 >109903493 >109903524 >109903664 >109905524 >109905566--Logs:>109906224--Gemma, Miku (free space):>109903299 >109904311 >109905177 >109905641 >109905840 >109906256 >109906954►Recent Highlight Posts from the Previous Thread: >>109904679Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109908812You don't just bring her with you?? Absolute monster...
>>109907388>>109907521>>109907546> saar 512 and 64k context are literally the same> saar one model outputting noise outside of three specific tasks and one model performing good in a multitude of different tasks are literally the same> saar they are the same why stoled saar
>>109907506Honestly ROCm on RDNA3 is so good I bought a 7900XTX to go with my old 7600 and everything is copacetic. Idc if nVidia is 50% faster or whatever, this shit is comfy and zero maintenance. It was unusable when RDNA3 first came out though.
release window for Gemma5?
>>109909190two more weeks
is gemma 4 still the best japanese translation model?
>>109909175It would be nice if it worked correctly with RDNA3 iGPUs like 780m, it works about the same on the RDNA2 680m I have as the 9060xt, vulkan faster than rocm for text, image gen works but slow (expected for the igpu but very disappointing showing from the 9060).>>109909226I can't really speak for japanese but qwen 3.8 has been doing a much better job at korean to english than gemmy.
Are current LLMs good with vision? I want to use Muse Glimmer or Gemma 4 in my LoRA baking workflow and have them judge outputs.
>>109909234They're alright, its definitely worth trying these days and not hard to setup
>>109909148I have no idea how to securely talk to gemma remotely
>>109909234Just have them judge outputs and you judge the models.
>>109909234There was an anon here who used models to sort and rate his images of women. It seemed to work pretty well.
>>109909257wireguard, tailscale+headscale, tuntox, tor hidden service, you have lots of options those are just a few
>>109909234They have no idea what good or bad is, just what's there. If you're specifically looking for features and want to ensure they're present then they're both fine, but they can't rank well.
what do yall think of thisqwen3.8 27 finetuned for creative writinghuggingface.co/Altworld/Hemmingway-1
>>109908194She looks like Cathe Blanchett.Nice ultra realistic.
>>109909314>i cant do thatits shit
>>109907469every time I try out newer models now I realize why I want to stick with dipsy since it didn't get pozzed like the others
>>109909286I have a hard time trusting Qwen with any creative tasks. Not sure a finetune can fix it.
>>109908812That's funny, I had the same feeling earlier. The masculine urge to share your adventures with a trusty robot companion, I guess.
>>109909314isn't the first one wrong
>>109909356In my day, people would just get dogs.
>>109909286The prose is legitimately solid but>>>>>QwenIt's fucking retarded at anything that requires any world or contextual knowledge
>>109909358someone flunked out of elementary school arithmetic
>>109909358It is wrong.>>109909392Don't try to confuse him. That's cruel.
>>109909405you're no fun
>>109909385What are you, two hundred years old?
did jev save local?
>>109909255>>109909262>>109909265>>109909280Thanks, I'll give it a spin! I remember reading here that Glimmer had better image understanding than Gemma, is that still true? Do I need different args in llama.cpp like more image tokens?
>>109909314I thought OpenAI was the best at maths. Why didn't openrouter check this ad before officially sharing it? How many retards were involved here jfc
>>109908134Damn elf seducing human men.
>>109909496For gemma:--image-min-tokens N --image-max-tokens NThe documented values for N are 70, 140, 280, 560, and 1120.I don't know about glimmer.
--image-min-tokens N --image-max-tokens N
>>109909496>Glimmer had better image understanding than Gemma, is that still true?Yes. 31B is still good so if you prefer talking to it over Glimmer you might as well use it instead. >Do I need different args in llama.cpp like more image tokens?Yes. Set both min and max to 1120. Increase b and ub to at least 2048.
CoT: Chain of Tards
>>109909523>>109909519Thanks once again, anons. I shall test both of them.
>>109909114
>>109909257I use wireguard personally but like the other anon said there are a lot of good options now.
I had no idea laya was the local version of jev.I'm getting it right now and it's so tiny it's only 400 Million parameters, finally something for poor people!!!Poor people rejoice
>>109907442Qwen3.8 27b's dialogue is pretty stilted even with finetunes that attempt to fix it.
>>109909572Spoiler: it's shit
>>109909385LLMs are better than dogs and bitches
>>109907332is big tech giving these guys huge payouts to quit?
>>109909581Where Jev currently appears better is decision quality on harder situations. On an independent identical-input benchmark, Jev beat Laya on triage (88.8% vs 80.0%), guardrails (96.7% vs 88.3%), moderation (98.9% vs 83.3%), and 12-way Banking77 classification (90.6% vs 80.2%). Laya actually beat Jev on AG News and MNLI.A larger 10,000-decision test showed a similar pattern: Jev 79.3% macro accuracy vs Laya 73.2%, with Jev particularly ahead when there were many possible choices. Laya was dramatically faster locally.
>>109909621Banking is an interesting benchmark.
>text-to-text is a trillion times better than it was two years ago>text-to-image barely improvedWhy is that?
>>109909387not even the 122b or flash next?
>>109907332he's a safety altruist jev
>>109909640Whats your issue with t2i our current models are some of the best we've ever had not even including text to video
>>109909640>text-to-image barely improved20-50 foward passes for an entire grid of pixels vs an entire forward pass for a single text token
>my name jev
>>109909640>text-to-text is a trillion times better than it was two years ago>text-to-image is a trillion times better than it was two years ago>text-to-video is a trillion times better than it was two years ago
>>109909621Just try it, it was made by a jeet btw
I miss the time when this general was about LOCAL models...
Using jev for speculative decoding.