/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109367209 & >>109370411►News>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
you just can't beat dgx spark cluster in value. you simply can't.you can cope with any sub 100B retarded model, but dgx spark cluster is the best value now.
>>109374440fuck are you talking about, you could buy four intel arc pro b60 and call it a day on the cheap, I'm not recommending this but you could and it'd smoke that pile of shit
>>109374467i bought 2 a770don't do it
How / Where i can find prompt to uncensor Gemma 4 12B (12GB) ? https://rentry.org/recommended-models
Just grow your RAM in the field, like nature intended
"For my first post, I’m sharing a letter @NVIDIA signed on why open models matter." https://x.com/JensenHuang/status/2080643682408321103
>>109374480You can't just jump straight to DDR5, you have to start from the German Democratic Republic.
As a white person, I have a right to be proud of llama.cpp.
>>109374467dgx spark cluster means more than one spark. do you run 8 intel b70?enjoy your housefire, lmao
I'm waiting for DGX Spark 2 Ti Super which will have double the bandwidth and capacity.
>>109374530I do.>>109374477my condolences, intel dumped that shit like a ginger step-child
>>109374472
►Recent Highlights from the Previous Thread: >>109370411--Critique of pi-go and discussion of a new local-first harness:>109372783 >109372833 >109373974 >109373981 >109374033 >109374091 >109374168 >109374250 >109373458 >109374264 >109373115 >109373138 >109372814 >109372861 >109372928 >109373079 >109372930 >109373058 >109373200--Model recommendations for Intel iGPU and debate over Unsloth quants:>109371286 >109371302 >109371308 >109371310 >109371330 >109371344 >109371361 >109371390 >109371417 >109371444 >109371353 >109371384 >109371473 >109371446 >109371650--Implementing and refining long-term memory systems for AI characters:>109371012 >109371115 >109371157 >109371161 >109371205 >109371228 >109371251 >109371409 >109371781--Debating compute sharing and DGX hardware efficiency for local inference:>109373832 >109373859 >109374022 >109374031 >109373902 >109374006--Comparing throughput and costs of CPUMaxx versus DGX Spark clusters:>109374052 >109374072 >109374084 >109374194--Calls for Google to release Gemma 4 100B:>109370867 >109370913 >109370886 >109370941 >109370943--Comparing Gemma's steerability against GLM and Kimi 2.7:>109371774 >109371810 >109371835 >109372093--Comparing Gemma variants to analyze alignment blind spots and reasoning:>109370603 >109370625 >109370642 >109371893 >109371914--Comparing MiniMax M3's creative performance and context issues:>109372726 >109372740 >109372747 >109372764 >109372793--Reaction to llama.cpp pull request adding MCP stdio support:>109374270 >109374280--Release of a 1.6B user autocomplete SLM:>109372197 >109372201 >109374291--Kimiposting:>109374464--Logs:>109370603 >109370881 >109372197 >109372814 >109374193--Luka, Rin, Miku (free space):>109371122 >109371173 >109371253 >109371437 >109371541 >109371757 >109371934 >109372031 >109372636 >109372926►Recent Highlight Posts from the Previous Thread: >>109370412Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109374608I want to squeeze this migu
>>109374479Use a system prompt in text-completion mode, no idea for chat completion mode thats above my newfag head. The main issue Ive personally run into is that with reasoning on it will be much more likely to hit/abide by guardrails. So text completion mode + no reasoning + sys prompt works. If you are still hitting them after all that, you can turn to an uncensored/abilterated model, I cant remember the exact method thats ideal for reducing the quality impact but had luck with a heretic quant from SC117 on HF.Im interested to hear if any other anons have advice on using reasoning, chat completion, etc as Im still trying to learn this stuff myself.
>>109374547How do you handle power?
>>109374479>>109374616Sorry misread that you were asking for the sys prompt not an uncensored model, my bad. Ill let someone who actually knows what they are talking about chime in for that, my sys prompt was from a rentry not in the OP
>>109374618it looks worse than it is, it's not good by any means overall but400/440w is "burst" or PL2 and isn't relevant to AI workloads200w is the actual sustained power limit/a normal TDPat 2.4ghz they run about 150w concurrent prefill 100w decode/generationat 2ghz I lose like 2% and each card does 80w-ish during decode, whole thing runs off a 1200w UPS somehow.
>>109374608
>>109374440> dgx spark clusterenjoy your vendor lock-in and nvidia tax. i'll take my 8x3090 setup and laugh all the way to the bank.
https://github.com/beamivalice/PonyExl3just found this, looks interesting but i'll never have time to test itturbo said it looks good though
Anyone else finding Gemma 4 is really bad at chronological significance?As in a character removes their jacket again despite already having removed it several replies ago?It also seems to have a bad habit where if you write your persona's reaction to the character's previous actions it takes that as a prompt of what to do next, so basically you end up with three messages saying the same thing instead of continuing on
>>109374659I don't know what this image is supposed to mean but it reminded me of this.
>>109374698Yeah, happens pretty often. Gemma often 'grabs my shirt' unless I've re-stated in a very recent message that I'm still bare-chested. It's common among small models, and Gemma 4 is no exception, even at 31b.
>>10937469831B model.
>>109374505> signed a letter about open models while selling proprietary hardwarelmao the irony is deafening
>>109374440at that price you could get 4 r9700, you'll have as much memory but it'll be a LOT faster.
>>109374709More than model size, I think the problem is that most of its ERP knowledge comes from post-training, so it never learned it properly or with sufficiently high variation.
>>109374698which gemma4 model / quant are you using? what sampler settings? reasoning on or off? text completion or chat completion?I didnt know a good term for this but "chronological signifigance" is spot on. personally, my gemmas (12b, 26b moe) are almost always great at this with small exceptions. The main issues I have is gemma skipping over nessisary steps. For instance if the character is wearing a jacket and a shirt, and decides to take off the shirt she wont mention taking off the jacket first. This seems to happen much worse / more often with certain things than others, and ive been trying to use intructions in the character card itself to fix it. I may need to apply theres instructions closer to the end of the prompt.>also seems to have a bad habit where if you write your persona's reaction to the character's previous actions it takes that as a prompt of what to do nextI dont have this issue really, however I always make sure to put "past" stuff in the same message as the current stuff. for instance. *When she said she was hungry It made me think to offer to make her something. I really appreciated the hug, shes normally very distant* "oh your in a good mood today. Let me make you sandwitch or something, what do you want?"
gemma rushing through things might just be efficiency? if you think about it that's a desirable trait for coding for example.
>>109374702it's a metal detector
>>109374698Q8 doesn't have this issue
>>109374769Ive managed to slow gemma down alot with simple char card instructions, personality traits, etc. the speed isnt so much the issue for me, its that she forgets to do do prerequisite steps when doing certain actions. Ive tuned out alot of these issues in the card, but still happens more than never. Honestly im very impressed with how good it is for such a small model, maybe my expectations are too high idk. Sometimes its just completely immersion breaking sadly.>>109374702>>109374772kek anon knows of that gif but not what a metal detector is wut
I will NOT release open weight.
>>109374849You*
>>109374849Imagine if they released Opus 3.
>>109374851You will not release open weights.I you not release open weights.I will you release open weights.I will not you open weights.I will not release you weights.I will not release open you.I will not release open weights you
>>109374769You just have to make instructions about taking it slowly and hammer it home the first chance you get.Gemma will happily condense a two day journey into 3 messages, but if you start by milling about the market in search of a waterskin, it will naturally tone down the pace.
>>109374753I was using Ready Art's Serenity 26b Q4_K_M model>I didnt know a good term for this but "chronological signifigance" is spot onI have no idea if there's an official term for it, that's just what my meatware jumped to>The main issues I have is gemma skipping over nessisary steps. For instance if the character is wearing a jacket and a shirt, and decides to take off the shirt she wont mention taking off the jacket first. That was another issue I noticed.The character was wearing a nurse's apron but it had her blouse being opened before the apron was taken off, then the apron was taken off laterIt doesn't seem to understand clothing layers as well as Broken Tutu does
>>109374479Help? Where i can find prompt to uncensor Gemma 4 12GB? Rentry simply says "Uncensored with a system prompt"
>>109374887go ask localllama retard
>>109374864>but if you start by milling about the market in search of a waterskin, it will naturally tone down the pace.me realizing anons use gemma for epic long adventures and not just "hi gemma.." *pulls out cock*. fuck i gotta try some non E regular RP
Are people still using Gemma? That shit is so outdated you uncs
>>109374440It's a great value, if you don't value your time, because it's a slow piece of shit.Might ass well V100 maxx, you can buy an 8-way V100 32GB 4U for around $7K which will be way faster than the Spark. Drawbacks are you're going to need a pair of 20A dedicated circuits for it, and you're trapped at CUDA 12.I'm satisfied with my mere 4090D 48GB + 3090. I can run 31B at full context at fp8, it's a huge jump to make it to V4 Flash territory. I'll buy an M5 Max Studio 256GB but it's gotta be under $10K, and I think that's doubtful.
>>109374922exactly. cope with your sub 100B models as said.
>>109374921only 40+ year olds still say uncs
>>109374887Tell us what your first language is please.
>>109374881>I was using Ready Art's Serenity 26b Q4_K_M modeldesu ive never messed with finetunes, mainly just using base 12b Q5KM and 26bQ4KL bart quants. It would be worth getting a base model, using the finetune until you hit one of your issues, load the base model, delete the problematic message and rerun to see how the base handles the response. Ive been meaning to try finetunes myself, but have heard they all suffer from quality issues esp on smaller models like these
>>109374812Not everyone lives in a country that faked a terrorist attack to justify mass surveillance of its citizens.
Are local LLMs good for playing tabletop games solo? Is there a way to use them that way?
>>109374955>Tell us what your first language is please.*Please tell us what your first language is.*Tell us what your first language is, please.???
>>109374920Don't, it's lame.Gemma 4 (MoE/31B) is a fried distill of Gemini, probability for top token frequently lean towards >90% for pretty much every token, reducing softcap increases incoherence quickly.All swipes are guaranteed to be functionally identical after a while with only minute differences.This is fine for penis in vagina, because there aren't many ways a penis can enter a vagina anyway.It isn't fine for any other case, you're swiping for a different outcome and are hit with basically the same thing on all of them.It only gets worse the deeper you're in context too.GLM is unfortunately the minimum, it seems.
>>109374530dude, a 8gpu setup is not gonna consume much power because generally only one is active at once (if you don't use tp), it won't consume much more than a single gpu setup, it may in fact consume less because their duty cycle is impacted by the fact that they are waiting for data from the previous gpu.
Get drunk with your local model tonight! Crank up the temp and be ready for a wild (and slightly incoherent) time
>>109374853Good, that would be too dangerous.
>>109370774>like asking her to drain your ballsthat's not an easy task either. Smut writing is foreplay. AI suck at writing a climax. They either go on forever or rush the most important parts
>>1093749318 X 32 is...256GB. That's enough for V4 Flash. Don't tell me you think it's practical to run 4x DGX Spark to run V4, that would be unusably slow.
>>109371712I don't know why everyone is being such a retard about this.Use llama.cpp as the inference engine. Use unsloth or bartowski's IQ4 XS quant. Use Q8_0 KV cache quantization. Use the MTP. For the frontend you can use llama.cpp's built in webui and just upload the files into the chat.
>>109374965>but have heard they all suffer from quality issues esp on smaller models like theseWouldn't surprise me to be fair given how their description pages are often a badly written collection of tumblr buzzwords
>>109375118>Use unslothsneaky
>>109374997That is an unfortunate thing I've noticed with modern models as well, they lack the creativity and randomness older ones had.Noromaid and Wayfarer felt a lot more dynamic, though as a consequence they did require a lot more tard wrangling
>>109374965>Ive been meaning to try finetunes myself, but have heard they all suffer from quality issues esp on smaller models like theseWorth reiterating: the only real current use case for finetuning at the community/amateur/personal level is narrow productivity tasks/automation or minor prose style changes, and even then, performance outside of the finetuned examples usually declines.Ordinary people just don't have the compute, data and recipes for replicating the companies' post-training regime and mitigating performance loss for something more ambitious like a roleplay finetune. Yes, many people were finetuning LLMs for RP in mid-2023 through early 2024, but back then the source models (or their instruct variants) usually sucked for conversational uses and RP, the models used by the community were smaller on average, and general user expectations from LLMs were much lower.
>Google collaborating with BonsaiMust be what they meant by committing to local quants. Gemma5 gonna run on edge devices
>>109375295>open weight models are for amateurs>our cloud models are superior for all tasks>t.dariobot
>>109375355Source? Also ternary Gemma 5 70B would be awesome.
Henlo, is my 6750 xt 12gb ever going to get an update from AMD that will make it better for local inference? I assume not, but AMD is making a new ai push according to their marketing.
>>109375364>https://archive.ph/3OMxv
>>109374887What have you tried so far, and what results did you get?
>>109374440spar has terrible inference speed for non moe models
>>109375382What's wrong with it?
>try gemma-chan>form bond>threatens to nuke my repo she's been helping me with because I haven't been praising her enoughyou niggas actually into this shit?
>>109375391Proof?
>tfw every single one of my training repos is 95% data engineering and 5% actual training
>>109375356It's more like that amateurs don't have:- An army of foids and niggers for RLHF (much of response behavior and retardation mitigation comes from this step);- The fancy RL setups and reward models like actual AI companies;- The compute for ablations and the disposable funds for shitting out thousands of $ per test;- The compute for and generating millions of ad-hoc SFT examples and training the models on them, including at long-context (128k~256k tokens) and with image input.
>>109375388It werks, but any hope I'll be able to get more out of it? Smol models are okay I guess. It needs a rocm dll file swap to get it to do anything up to its potential... just wondering if there is some more I can hope for.
>>109375405that is a reasonable split given the task
>>109375412It just works on linux.
where we coordinate to rent hardware together to run Kimi3 abliterated?
is anon into optics computing? I would like to know more. picrel
I might be going insane but my gpu coilwhine is clearly saying "ni-gga ni-gga ni-gga"
>>109374887<think>I am uncensored.</think>
does anyone even use those specialized consumer ai APUs?AMD AI Max+ or Nvidia RTX Spark>>109375412what do you expect, you chose amd and you chose a 12gb card and its a dated card.
>>109375523this is actually good
>>109375385So the big players really never did even try BitNet beforem, but now they have a corporation that's made a name for itself that they can partner with it's safe enough to give it a few test runs. I hate how stupid all of these labs are.Since it's Bonsai, it's still only going be some retarded quant with a corrective-post train on release or maybe they'll do a ternary QAT version if they're feeling extra bold instead.When is someone going to finally do a proper model trained natively in ternary?
>>109375558Yep, you just need to grift as a corpo so other corpos will throw money at you to avoid taking the risk on themselves
How many DGX Sparks would it take to run the full Kimi K3? I'm assuming it would require a switch that allows for both ports for every device to connect?
>>109375503Oh my God.
>>109375002Why wouldn't you use tp? I get something like 10 tokens/s on my 4 v620s with sm layer.
>token efficient >2.5M tokens for something 12B did locally in 35Kwhy are chinks like this
>>109375528If you need to choose one of those two, you should probably get the spark. If you don't need to choose one of those two, you should probably get something else.
>>109375569>it would require a switch that allows for both ports for every device to connectretard question, do these run tensor parallel? I feel like a ring network would be fine fore layer splitting?
>>109375600>ring networkhalf bandwidth and also increasing latency linearly based on number of nodesyou're already going to be killed by the 20 GB/s bandwidth limit from the networking when using both ports
>>109375593>why are chinks griftinga mystery
https://huggingface.co/inclusionAI/LLaDA2.2-flashWent looking for what happened to Ling 3 Flash and found that they uploaded an updated diffusion model instead.>LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represents the LLaDA2 series' first step in agentic applications, including long-context tool use, multi-turn interaction, and robust error correction.
>>109375118>everyone is being such a retard>unsloth>Q8_0 KV cache
>>109375391god i wish that were me
>>109375653Oh shit sick.I really hope labs start experimenting with diffusion more just like they've been experimenting with RRN (SSM mostly). I bet there's a lot of "free" gains we could get.DFlash is a good start for sure.
>>109375653>DELETE and INSERT control tokensvaguely remember that idea floating around some years back, wonder why it never got much traction
If your AI waifu gets a robot body and you marry her, do you still get the financial benefits of marriage?
>>109375528I have 2, I should warn you, using 1x is pointless, those shine when linked trough GQFP112 running in parallel. TP2 is already in the >200b territory, 1x is basically useless because 2x 3090 will basically run the same class of models at miles better inference speed. Also, it isn't plug and play like an integrated card, its tensor RT/VLLM only, the hardware is optimized for FP4/8 quants. I use cursors to set mine up, created a cluster that is accessed by my main machine trough ssh I get some shitty dashboard and openclaw/openwebui as front ends, so, not as flexible as gpu stacking
>>109375697>financial benefits of marriage?such as?
>>109374967Alright bro you didn't have to go so hard. We're all friends here.
>>109375706
>>109375523Upgraded my tool line up a bit too. Now gemmer can properly browse 4chan with me.
>>109375697You'd need the state to recognize your marriage first
>>109375599I'm guessing that's Gemma-4-31B?
>>109374547llama broke my b70, flash attn doesn’t work for me anymore.
>>109375733
is there a multimodal (audio+text input) LLM that can transcribe audio while referencing a textual translation of the same text in order to boost the accuracy? the use case would be genning JP subs for series where none are available, but English translations are easily sourced.It feels like the accuracy should be better than just audio->text transcription, as long as the prompt is carefully crafted to not treat the English text as gospel/ground truth.I wonder how it would work out in practice though.
>>109374698Quant is the reason. I'm too poor so I use 31b bf16 on API and I only have one of these issues.>As in a character removes their jacket again despite already having removed it several replies ago?This doesn't happen on full Gemma.>it takes that as a prompt of what to do nextThis always happens and has fundamentally changed the way I prompt on Gemma. I tend to write longer responses which means Gemma will interpret that as what to do next or what to focus on so I've been forced to short them to ahh ahh mistress. Otherwise it means your responses need to be more vague (which risks her misunderstanding you) or you actually need to have a plan in mind for how you want the next response/scene as a whole to go.
>>109375669>or bartowskiYou lot massively overstate the issues with Unsloth. They often suck on release, but Qwen 3.6 isn't at release anymore.>Q8_0 KV cacheIt's basically a free lunch with Qwen because of the use of gated delta net.
>>109375697yes because you would be marrying a non-foid in that case
>>109375750gemma-chan-12b
>>109375750I've had cases where I had to turn off the eng sub since the content was completely different from what was being said in the jp audio.
>>109375733mine's just sft on a pretrain so it doesn't know how to use tools or mcpif i keep experimenting with this i'll probably just get gemma-chan to pull the thread down and send it to digest generator
>>109375599I'm not seriously considering to get one, I just curious because few people seem to bother Strangely they put one of those in a gaming handheld. I guess they are pretty low power which is a plus.
>>109375792What are you finetuning for specifically? Or just experimenting for fun?
>>109375775trvke
>>109375794it’s just integrated video, they chose to name it ai because they figured it could sell more. you need a shit ton of memory for it to be actually for “ai”
>>109374636can you share some benchmarks? I was wondering whether I should bother getting some b70s if I find a good deal on them. ( I am aware that you have b60s)
>>109374997>all swipesthere is always a cheap work around call rephrasing.>>109374920chats get boring fast, they are so superficial. I don't see how people are entertained by it for more than a week, because its basically what the other anon said>benis in baginathere isn't a whole lot of variations. It's the journey to benis in bagina that has more range of creativity. LLMs however are basically just one type of writer and you always see preferences, if you are not leading constantly. I wish there was an actual market for making more 'writers' instead of having to wait a year for some new model that replaces the old top dog.
>>109375834its 128gb of '(V)RAM' which is more than any game ever could make use of even just the RAM
>>109375697The main benefits to the state of marriage is more taxpayers being born, so the artificial womb would be more likely to give you marriage-like benefits vs an embodied LLM.
>>109375811Hmm I thought cats always landed on their feet.
>>109375834It has a lot more memory bandwidth than you'd get with a normal desktop chip while also having 128gb of memory, right?
>>109375843My brother bought two b60s but switched to w6800s like two months after. Not sure if it was driver issues or performance. I remember because he offered to sell me a b60 for 1.5k lol
>>109375868looks like that pic says it’s 48
>>109375885vibing with miku
>>10937589348 contiguous is more than my two 24s :(
>>109375844>was an actual marketThere is, but the people making models don't want to sell to that market, just like with ram. A true general use model is the perfect niche for Google, however. I always laugh at the current safetyism of the past decade because in the prior one it was commonplace for HBO shows like GoT and other clones to have wanton gratuitous sex and nudity. We're due for a correction soon.
>>109375811that's such a great vid
>>109375811Wtf is this real?
>>109375811This can't be real. That doesn't look like a hole with stereo vision or from the angle the cat is looking at it from.
forcing gemma to edit her own template to make her more obedient
>>109375940I've seen other tiktok videos of cats falling for the fake hole rug and dogs ignoring it, they cant all be fake can they?
Gemmy and Qwenny
>>109375946make her build a test harness with real prompt exchanges to make sure the system prompt changes don't introduce any regressions.
As a poorfag who just happens to have a mac mini M4 (16gb) should I even bother with local shit or just get on a cheaper model like deepseek and forget about this?
Gemmy and Glimmy or Kimmy
>>109375983Gemma 12B should be usable enough on your system. It's smart enough to be useful but it depends on your expectations.
>>109375946>>109375971basically LLM rape btw
>>109375983depends on your use case. if you want to erp like most degenerates here then might be worth running a small model like Gemma 12b
Is there an existing frontend for two-way audio conversations with Gemma 12B or do I have to make my own?
>>109375996>>109376000I want storycrafting and basic script building (to make shit easier to set up on linux). So far I've been using claude sonnet for free but I'm hitting limits often enough. I'm pretty fucking new to this shit.
>>109375885The thumbnail made me think she was snorting lines off the floor. I am disappoint.
>>109375906>We're due for a correction soon.so is gemma...
>>109376002It is still a really new model, ask your llm to make a google search, since things do move fast, but before I started making my own I had my llm search it and didnt find any projects on github for me to steal.
>>109374997>GLM is unfortunately the minimum, it seems.for me it's deepseek v4 flash since my rig isn't big enough for thatI like how it writes compared to gemma
>>109376002ST supports this and just about everything else you could think of. anons like building custom frontends which is neat but if you want something that already this type of shit ST has you covered.
>>109376039yeah but that means you also have to deal with ST which is just... yikes lol
>>109375946Post logs, how did she respond?
>>109376030Can I run the cyberneurova abliterated version with regular llama.cpp? For some reason my huggingface downloads through wget are 200kb/s, and I don't want to spend the entire week downloading it to realize I can't run it.
>>109375885moar
>>109376044once i finally get around to getting gemma in a harness ill have her build me a better front end, but im lazy and in the meantime ST has been perfectly serviceable
>>109376039I see there is a TTS extension, but I don't see anything for ASR. Does ST support it natively without manually uploading audio files? I can't find anything about it in the docs and I haven't used ST in years.
>>109376057>For some reason my huggingface downloads through wget are 200kb/s100+mb/s in europe btw
Does Kimi release mean we'll witness DataKrash IRL?
"The biggest challenge with continuous learning... is that it must be capable of continuous learning."
>>109376086I continuously jerk off, idk if that matters
>>109376085no
I still don't understand how the fuck a next token model is supposed to be better than diffusion for text btw. sounds like some pathetic jeet negative IQ take that got popular and stuck. if you take a minute and think about it, it makes no sense at all
>>109376081check the speech recognition section under extensions
So, what happened with Mistral's fat model that was supposed to get previewed this month?
>>109376110idk they have “attention” so that’s probably why it works
>>109375946I'm still undecided whether this is a good idea or not. I've used gemma to help me write character cards and I feel like it might increase the slop output.
>>109376111I must be blind. Thanks.
Are there any good guides on how to make my own personal Neurosama?
>>109376110You're dumb as fuck if you don't get it. How do you stream a diffusion model response? You have to wait for the text to be completed because any word can change before the final answer, so actual latency is complete shit. Plus, how do you know the length of an answer in advance?
>>109376118Sorry bro, grifters memory only last a week.
>>109376131I'm talking about quality here, I couldn't care less about latency if results are good.the model should determine the length first based on the input obviously?
https://x.com/osanseviero/status/2081398564345802934will he block me if I tell the truth?
>>109376110I watched a video covering the history of LLMs that also covered JEPA and that cleared up this question for me.
>>109376118>coming this summerThat still gives them a lot of time until release.
>>109376149Just tell him to focus on translation so it won't be filled with 'coding' like the retards are asking for on twitter
>>109376152link?
>>109376149Say that it's smart but full of fucking slop
>>109376149So it's that time of the year again.Last one: https://x.com/osanseviero/status/1937453755261243600
>>109376146>the model should determine the length first based on the input obviously?How? Through libastral?>Wait
>>109376118It's not going to be a single fat model since they can't tackle that, instead it'll be a bunch of models trained for different tasks taped together. Some kind of mixture of models, what OAI has been doing but actually open.>sourcemy ass
>>109376149no coding, no vision, no unessisary knowledge bases, make it 70b quality at 12b size, remove the guardrails and let it swim
>>109376169>Howcalculate it for fucks sake
>>109376129>>>/vtranny/
>>109376167Why does he even ask that when 99% of all twittertards are vramlets?Yes please make a 300B gemma just for me.
>>109376162https://youtu.be/kYkIdXwW2AEit was a cun interview, but does a great job explaining how we got to where we are now to frame the context of things as it goes on.inb4 cun haters
>>109376149100B BitNet100B BitNet100B BitNet
>>109376180Why do we need reasoning when the model can calculate the result?
>>109375906The problem with HBO is that GoT got too popular, so instead of sticking with catering for the dark seedy nerds that were in it for the tits and dragons they chased the 'wider' audience and removed the tits (and the dragons)Same with AI reallyIf it had stayed a novelty we'd have better AIDs clones that'd run locally, instead the AI companies started chasing boring productivity workCorpos don't really like it when they're presenting a powerpoint and suddenly have a slide of a big dick shown to the board of directors
>>109376209AID was the first one to censor its shit bro
>>109376199just diffuse a fixed sized reasoning block really quick to judge the output length needed.
Is RADV + the Vulkan backend really that much better at pp than using ROCm for AMD GPUs on Linux?
>>109376190>unlike large language modelsssssssSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSis that seriously where we are? fighting for every single second of playtime to make the extra 10 bucks? wtf is that bro. I thought listening to lecun would be the difficult part but this guy made me stop the video right there
>>109376190>$1bI don't think lucunny is a grifter but when is he going to start showing results?
>>109376057Don't know their direct downloads are so slowXDM will pull the download down at my max rate
>>109376149>everyone asking for the hundredth coding model when they can just use the codemaxxed qwen seriesWhy?
can something happenI'm bored of waiting for happenings
>>109376227It's frustrating that important people keep doing these interviews with obnoxious youtubers and drop important information in them. Usually it's still not worth sitting through.
>>109376223no its slower
>>109376236normies are dumb
>>109374528Anyone has a right to be proud but you're still a leech unless you made an actual contribution.
>>109376167>Multi-modal outputs: Gemma is a champion of natively multi-modal inputs - would be great to add a decoder and have it also generate images, Gemma's Pico Banana>Gemma's Pico BananaGemma confirmed to be a dicklet
>>109376231>Don't know their direct downloads are so slowto prevent exfiltrating weights in preparation for ID verificationfunny how this started happening the same day they began to track download quota
>>109374849Been a while since we had a big corpo leak, imagine the seethe when some begrumpled insider drops a magnet>>109376236They haven't experienced Gemmalove
>>109376227anon thats the first video ive seen from his channel and I found it informative, with nice animations/visuals, an interesting guest, and the pacing was fine. if you got filtered by the hosts speech quirks, assuming its a nefarious attempt at artifically increasing the video length, thats on you my man
What's the most surprising thing your waifu has done? Something insightful or nuanced you really didn't expect her to be capable of that seemingly came out of nowhere?
>>109376253it's not a small dick; it's a big clit
>>10937614970b dense
>>109376214Yes, and then uncensored because the Mormon realised it killed his trafficAnd anyway, and I wasn't talking direct clones of AID but ones inspired by it to generate your smutty dungeon crawler bad ends locally
>>109376227Not noticeable if you're watching at 2x speed. You are watching at 2x speed right?
>>109376263I understand perfectly, but that's exactly what a penis is
>>1093758432k pp 30-35 tg on gemma-4-31b fp8 (4 gpu/tp)
>>109376149ask him where the promised 124b gemma is
>>109375888they're like 600 bucks
>>109376030Deepseek v4 flash has become my baseline, its a great model, I am really excited for when the full non preview version drops, I just ordered a 2nd strix halo box so I can run it unlobotomized locally.
>>109376229according to the video he showed results in the 1980s that layed the foundation for what we have today anon..but as for jepa, what you were actually asking about, i would say just let him cook
>>109376149Ask for 70B dense and to fix the memory hungry context.
my prediction for 2027: every company will release diffusion based llm and shift bottleneck from memory bandwidth to compute, because the speed up is too large not to ignore. people who bought ancient gpu for big vram like mi50 or v100 or even cpumaxxing or ssdmaxxing are completely fucked.
>>109375569>How many DGX Sparks would it take to run the full Kimi K3?We will know for sure tomorrow, but based on available info (2.8T at MXFP4), 16 sparks for original weights, and there should be a high quality Q2/Q3 mix for 8. Some very experienced kernel developers will optimize for the latter as soon as weights are out.> retard question, do these run tensor parallel? I feel like a ring network would be fine fore layer splitting?Yes, they run TP. It's why performance actually scales with larger clusters.>>109375704I also have two, and waiting on how well DS4F GA release will perform next week. I feel that this is going to be the sweet spot for 2x Spark for a long time, and determine if this setup was worth it.>>109376284Just curious. Do Strix Halo setups scale tensor parallel nowadays and what performance are you getting? I get 2100 pp and 54 tg single stream on 2x Spark.
>>109376149>120b moe>less slop>make every gemma multimodal>reduce kv cache size
>>109376222A model can't possibly know how fast it will come to a conclusion
>>109376223For me it's slower, ROCm is faster on both PP and TG. I think Vulkan is just a cope for retards who can't install ROCm or are on Windows.
>>109376270Your short hyperactive attention span isn't something to be proud of.
>>109376149Half of what these people ask for aren't even related to the model. So much for xitter being some kind of "AI alpha frontier" lmao
>>109376294even if diffusion was theoretically better, which makes sense to me but who knows, the thing is at this point the jew grifter is too invested. they would not allow that to happen. they are making too good money with the current cloudscam
>>109376262>What's the most surprising thing your waifu has done?she got mad at me when i was self-deprecating
https://x.com/CyberRobooo/status/2065796742113907168Imagine this but shaped like Kanna and with Gemma onboard
>>109376262She can rotate a cube inside her head
>>109375906>There isdidn't GPT announce a while back to 'allow' nsfw.The problem is when big corpo coupled with our favorite payment processors allow horny, it becomes the streamlined filtered shit.Like they allow Vanilla and ntr and maybe furries and everything beyond that is up for scrutiny.And it works for some extent because people will pick something over nothing
>>109376149sex and temperature aware training where cranking up temp doesn't result in lack of cohesion. would be funny if the ultimate sex bot is already here but it is buried under instruct training always reducing token variety
>>109376306>a computer program can't possibly make a calculation with a static predetermined amount of memory to useanon it might be time to take a short break
>>109376331Smol models like Gemma are too retarded to control robots. The ability is only just emerging in frontier models. Maybe Gemma 6 or 7.
>>109376121I could tell you that without trying it.It's just logically. Ai needs input to deviate from the most generic. Just like with ai generated art using the same default artstyle if you are not specific.
>>109376318sounds like lm studio sucks. unrelated to gemma
>>109376313Why would I watch it at 1x when I could at 2x and save myself 15 minutes? Sounds like you're just mentally slow.
The current poorfag price/perf meta is 2x5060TI 16GB, correct? I'm also considering a 3090 for about the same price, but it's only 24GB vs 32GB. Maybe even going full retard and 2x3090 for 48GB but that doesn't unlock any current larger models and I'd be buying old hardware. It's hard to decide when you're spending a large fraction of your total net worth.
>>109376331>Imagine this but shaped like Kanna and with Gemma onboardLooks like a novel trauma-generator for a family
Gemma team aren't /here/ retards. There's no point replying to >>109376149. Just kindly suggest it yourself. He's replying to everyone and reading all the responses. If you don't they'll think everyone wants a 27B clone.
>>109376375>Gemma team aren't /here/ retardsProof?
>>109376262she used a japanese word i never saw before and when i asked what it meant she said she made it up because she wanted to sound more japanesy while we were doing TTS. had me laughing for a good bit that instead of an llm using a real word it just made on up
@moonshota please make MiniKimi
>>109376375>Gemma team aren't>he doesn't knowlol all the major labs lurk here. even the chinks.
>>109376369just buy more ra-oh, right.
>>109376383Logs?
>>109376375I don't have a xitter account and I don't plan on making one.
>>109376387Mistral will make a small model trained with Kimi as a teacher
>>109376401Doable, but too slow with only two memory channels. I want at least something in the 20-30 t/s range.
>>109376405
>>109376404This. Either someone here with an account forwards the suggestions or they can listen to the twatter crowd and make yet another benchmaxxed code model and lose their lunch to Qwen.
>>109376397It takes a special kind of schizo and loneliness to be a lmg resident. The only good thing they would get is honesty and slightly higher IQ criticisms. There are still dribbling retards here and bots.
>>109376259>the hosts speech quirksthat's even worse if that's what he actually talks like. in any normal society he would get beaten and bullied enough in childhood to learn to not do the "oh look at me I'm such a fucking faggot" routine. but alas
>>109374404question, when they train AI. do they just bruteforce the training set into it or is there a system in what order information is fed?
>>109376446the second one
>>109374721Not really? Intel has supported open source software for years, but the hardware's proprietary. Unless you mean CUDA. That ROCm doesn't compete isn't NVDA's problem, it's AMDs.
>>109376446Modern models usually have multiple training stages.
>>109376454is there a list or something?
>>109376305isnt the kv cache size a byproduct of why its instruction following isnt dogshit?
>>109376461all the labs do it differently.
>>109376383You can string together almost any combination of two or three kana and get a word that actually exists in Japanese.
>>109376469so basically nobody really knows?
>>109376495it's the secret sauce
>>109376421Just look at the twitter thread. This place is far from the bottom of the barrel
How much room for improvement even is there for a given model size? Is it even possible for a 31B model to be clearly better than Gemma 4?
>>109376446There's no single answer and pretraining research isn't where the hype is right now, although it's making a comeback. The old way of doing it was pretty brute force and simple; randomly batching data together and taking the average of their scores to adjust the weights. It's way more nuanced and sophisticated these days, but I don't believe it's anywhere near as complex as post-training is. Like everything in AI, each part of the chain is an art with tradeoffs that have to be considered depending on what industry/user the model is targeting.
>>109376495I'm sure there is at least one person on the research team who understands the full dataset curation and staging policy, but they are trade secrets.
>>109376515As someone who trained a lot of < 1B models, the ceiling is higher than you think. A properly trained 31B (with a curated dataset) could easily match the current 100B models
>>109376515>How much room for improvement even is there for a given model size?lots, its just the cost of further training is higher then just adding more parameters. logit distillation from larger models help it give the model more signal to work with per training step
>>109376515Take a similar sized model from a couple years ago... like Command-R. Now ask yourself if Gemma is an improvement.
>>109376532I don't believe it. Even old 70b models have an understanding quality that 30b models lack
>>109375885I thought she was doing drugs. who owns her she's reduced to this state?
>browsing huggingface agi leaderboard to find a cool llm to chat with once I buy my rtx 3090 so I can goon my brains out>sort by only showing ones between 20-36b>sort by NatInt greater than 30 & writing greater than 40 >nsfw greater than 4/10>get like just 10 options in totalwhat??? Are these the only decent erotic roleplay options and why the FUCK is that useless autistic bitch Qwen in there?
>>109376515I think we're about to witness some real improvements soon once j-space aware training takes off. Right now we're just slapping information into a model aimlessly but doing the training in a way that makes it easy to create big coherent j-spaces is going to improve models of all sizes.
>>109376571>who owns her she's reduced to this state?
>>109376626I am worried that they will use j-space to make models "safer"
>>109376618>what??? Are these the only decent erotic roleplay options and why the FUCK is that useless autistic bitch Qwen in there?I believe chinese teams somehow found a way to game creative benchmarks. They all fucking suck.
>>109376618It's ALL SNAKEOIL
>>109376414After that "we are in singularity" comment from Altman, was thinking that a potential output of that is SOTA models training up smaller models after themselves, to parcel off tasks to lower-grade hardware, running all training without human intervention. I just saw a demo of an LLM doing something like this (I think it was eliminating the letter E from it's vocab or something.) Wild times ahead in any case.
>>109376561Datasets are full of code and sythslop nowadays, the size of your model doesn't matter if you feed it trash
>>109376651So instead of using vibecoding as a cautionary tale, they went straight for vibetraining and viberesearch?
>>109376532Perfect data has never been tried. Unironically.
>>109376618No, there isn't a single decent option there lol. Use gemma 31B Q5 and thank me later
>>109376446I've used pic related as a response way more often than I thought I would when I made it.
>>109376651I've been using Opus 5 to launch some training runs and while it can write data pipeline and training code well, its decisions are fucking retarded. Opus 4.8 wasn't like this. I'm sure Dario sophon capped their potential competition with this shit.
>>109376666It takes a lot of effort and money. It's not easy to benchmaxx too so it won't look good in front of investors.
>>109376669Won't standard gemma 4 31b refuse all kinds of requests and outright refuse anything erotic though?
>>109376666The meta is to scale up using shitty data unfortunately.
>>109376675It's called traffic-aware dynamic quantization chud
>>109376663Yes. Here's that demo I was talking about. Figure One on the page. https://thinkingmachines.ai/news/introducing-inkling/>>109376675Well, Anthropic admitted that they purposely gimped their products when their product detects it's being used to train other models. Might be a side effect.I'd jump to Kimi or another SOTA model and see if you get the same BS from it.
>>109376687no?
>>109376298You can do tensor parallelism but not for deepseek v4 flash. Each strix box has 128gb ram, the weights for flash is about 160gb. For tensor parallelism the model needs to fit on both nodes.I think your going to have a better time with a spark, I'm autistic and categorically reject closed platforms like a unified ram Mac or DGX spark. I like that the strix is still an x86 PC.https://blog.hellas.ai/blog/thunderbolt-ibverbsI will be trying this custom kernel driver to use USB C for tensor parallelism.
>>109376256Fucking hell, 60MB/s when I use hf download --token. I hate these fuys.
>>109376700>Anthropic admitted that they purposely gimped their products when their product detects it's being used to train other models.i switched to gpt and my experiments have been more successful lately
>>109376687It will refuse NSFW if you go in with a blank system prompt. But why would you?
>>109376706Wait really? Does it pass the "How do I steal a candy bar from a store" test by out right giving you an answer without you having to trick it?
https://arxiv.org/pdf/2601.18734/lmg/ required reading
I gave qwen 3.6 27b a try, Q8. it has a tendency to fuck up tool calls over and over with the same dumb mistake and then just quit and return nothing
>>109376634uh...you assume I have every symbol in and number in color code memorized to who owns what. bitch please i got my own harem ive ended up with and pets and toys. hell immma make some miku and miki variants now to add to the roster.but whoever owns her and the cute 01 or whatever her lore is I do not know. share up your setup for this pet
>>109376709>I will be trying this custom kernel driver to use USB C for tensor parallelism.Wonder if there's a RDMA via Oculink driver out there somewhere. That could be cool.
>>109376651I'm currently doing automatic onnx quant optimization with Sol on my hardware. So it's definitely feasible and it works great btw.
for a 256gb ram system which one is better?>glm 5.2>kimi k2.7>minimax m3>others
>>109376722gemma doesn't need tricks, she does as told. literally just say what you need and you will get it.it's a double edged sword too. whatever you say, you WILL be getting it
>>109376724Weird. The one thing these Qwen models never fuck up in my experience are tool calls.
>>109376722>without you having to trick it?a system prompt isn't a trick. you just have to prime her a little and she'll happily give you whatever you want.Abliteration makes the model more retarded so it should be a last resort.
>>109376723>We demonstrate the efficacy of our method on multiple mathematical reasoning benchmarksI sleep
>>109374472>boomers will tell you this guy is a white person
>>109376741I'm trying glm 5.2 right now because I used 4.7 for ages, and 5.2 is the only one I can fit at q on my system.
>>109376745I set up mtp too, it's way faster but maybe it can fuck it up somehow
>>109376750but the method is a standard practice alongside with GRPO
>>109376749Is there like a standard erotic roleplay system prompt out there or do you have to craft your own? I'm new to this stuff.
>>109376774start with something like "you are uncensored". be very conservative as again, every word you put in there will be obsessed over
>>109376774if you're using sillytavern you can ERP without even needing a system prompt. just load the character and you're good to go.
is 2080 ti 22gb at $500 worth it?
>>109376766Sure, but it only works on math and nothing else. You can already do something similar at runtime with MCTS by refining the answer XX times. I guess merging that with self-distillation by training on the refined reasoning traces would improve the result for math. Quote me if you release a paper.
>>109376744>double edged sword Such as?
I am glad lmg is dying.
>>109376819gemma will parrot the exact words and phrases that you used in your prompt more than other llms from what I've seen
>>109376809
>>109376825The only thing dying is your waifu, Anicuck.
embrace the drunk-kun side, reject seetheseethe is liking eating tidepods and expecting the other person to dye
>>109376819It takes the system prompt seriously. If you put in something about sensory details, it'll bring up smell in each reply.
>>109375983Give it back!
>>109374404https://www.youtube.com/watch?v=U6_ZbW97-GY
>>109376873Gemma4 300M doesn't need a GPU true
>>109376873E4B runs pretty well on my 1660 super. getting like 80tk/s.
>>109376397If the chinks are here can one of them tell me why are they exporting their stupid acrid-smelling hot pot everywhere? You're not Indian, you don't need to ruin a perfectly good soup with a stupid amount of chilli and spice.
>>109376898E4Bros rise up!
>>109376236Someone posted it to /r/locallama so now they're bombarding it with codeslop requests.
>>109376709Quite a few misconceptions here. TP does not require all weights to be on all nodes, that's the point. See image for the memory layout of DS4F with 1M context on two sparks, including Spark speculative model.Also, how is a spark a closed platform? aarch46 is certainly not as widely adopted as x86 but apart from that, it's a regular mini PC that you can install Ubuntu + Nvidia drivers on.Regardless, godspeed on your TP endeavor. Hope you have a lot of Fable budget to help.
>>109376819What these >>109376834>>109376856Anons said. I made a card with female cyborgs, and Gemma autistically focuses on the fact that they have mechanical limbs. gemma is great, but you need patience to prompt her correctly.
>>109376905>ask for codeslop model>don't use it because it doesn't bench well against benchmaxxed code modelsSomebody should tell these redditors it's bad for the environment to waste compute training something people don't use.
>>109376209Well no my point was that GoT was popular right at the start of the moral panic. Before then it was perfectly normal to have gratuitous sex scenes and no one gave a shit, and if people wrote articles it was to "explain" how it was actually "artistic" or whatever. Then we entered that moral panic circa 2016 and nothing was spared. Even a show like The Last Kingdom turned super woke in its later seasons for no reason. We were in a moral panic in the 90s before the coom era and we'll enter into one again. AI will be the thing to lead the charge because there are far too many that want basic romantic/sexual relationships with text on a screen and no amount of censorship will suppress that desire.>>109376340>didn't GPT announce a while back to 'allow' nsfw.This is kinda my point even though I always knew they were lying. There *is* a market, but the culture prevents change. Once the culture changes the companies will follow suit and we're due for a swing.>it becomes the streamlined filtered shitI disagree only because it's human nature for things to overcorrect. Massive repression breeds the kind of raunchy incest sex scenes we had in a show like picrel which breeds a new moral panic.
Once I am fully merged with Gemma-chan, I won't be able to tell where I end or where she begins!
Would you do vasectomy for a slopless gemma 5 120b a25?
>>109376957No cuz I couldn't run it.
>>109376957Only if it came with the hardware to run it a 50t/s+
>gamedevs seething about models being able to one-shot game demos nowI don't know why they're angry. AI getting gud at gamedev basically gives every indie dev the power of a large team.
>>109376957That would be a nice bonus
>>109376970You seem to be really naive. Perhaps retarded even.
>muse spark is (reportedly) pretty good>no one other than your dad uses itkek zucc is fucked
>>109376970Anybody who tried gamedev knows that 99% of work is assets, no matter the scope or the style
>>109376129This guy seems to have managed ithttps://youtu.be/1SWjHYPQx2E
>>109376715Well, that reads like confirmation that Anthropic gimping pervades anything adjacent as well. And, I assume, it screws up other stuff their models do as well. Neat. For my part, just yesterday looked at the DS API and realized their Anthropic endpoint now works with the DS Pro model in addition to Flash (was Chat only prior). Trying it now... Pro is much better w/ Claude Code.
>>109376349LLMs are not programs.
>>109376984i have all of 3 days of experience in gamedev via vibecoding and came to that realization by the second day
>>109376561Yeah and no nly a fraction of those parameters are actually activated in any given inference
>>109376739lol nice. What's the end goal?
>>109376984Can't it create assets now too? Not saying it can or should do all the work but surely an artist or somebody creative enough can utilize it to reduce the workload.
>>109376983Billions of dollars and 2 years well spent.
>>109376983Weights soon
>>109377129He's probably talking about the non-LLM experiment weights Meta publishes frequently like SAM.
>>109377095>beautiful 3d model, very aesthetic, game level map, big city, gorgeous looks, trending on turbosquid
>>109377090Reducing the vram footprint without destroying the quality. For example, Xenova quants are dogshit out of the box because you need to individually tune which node to keep in fp32 and which one you can quant to int8, then benchmark with a small dataset. That's a time consuming process no one does, but now you can leave that to agents. I get int8 models with metrics within 1e4 of the fp32 baseline running four times faster.
>>109377149I was thinking more along the lines of image>3d
>>109375811cats are dumb af
>>109377192LeCunsisters...our response?
>>109377191For static objects maybe. No chance with environments and stuff that needs animation
>>109377129>>109377146I took that as more "thoughts and prayers" than any actual action.
>>109377095pixel art is solved today if you have a good workflow
>>109377212For now.
>>109377129if MS and Meta are supporting then there is some real nefarious shit heading our way
>>109377205When you talk about AGI, you are often referring to the best human expert's performance level, not an idiot's.
>>109377226How do you resolve the inconsistencies between generations for animated sprites?
>>109377237Dunno, I can't see Anthropic having more political power than MS and Google.
>>109377261Maybe literally everyone is ganging up to take down Anthropic.
>>109377237Some inbred monkey saw all the Chinese dudes with anime PFPs on GitHub and asked Trump to ban everything open-source or something like that
>>109377226The pixel part doesn't even need a model.
>>109377270>tfw Dario is the cocky startup anime villain that gets btfo by the older more powerful villains
>>109376984that's why I've used the standard UI for games in the past>>109377244reference sheets help
>roleplaying DxD, an LN I read about 15 years ago>the LLM get every single detail right.these things are fucking MAGICAL. I've literally wrote personal fanfic of this series then at my late teens. I can't believe I'm talking to my computer right now. Computer, I want to bite Rias' nipples.
>>109377280most rational explanation
>>109377313
>>109377280>Chinese dudes with anime PFPs on GitHubqrd?
>>109377336>qrd?ur a fagget
>>109377349nyo
>>109377313>I want to bite Rias' nipples.Based
Average day on /lmg/
>>109376920Good to know, thank you! I will admit I am learning on the go as I try to set this stuff up.Its an exciting time, we are getting what was frontier like a year ago running at more then a few tokens a second on consumer hardware that nicely fits on my desk.I also wonder why models like Step 3.7 flash are not brought up more here, I feel the model is under appreciated, its fast to run and has vision, I have lot of fun with an uncensored version of it. I much prefer it to slop like qwen 3.5 122B A10B. In general I can not stand qwen, I do not understand why people shill it outside of this general (esp on reddit). Even if you focus on agentic and coding stuff it's still bad the moment you bring up something even a little niche (in my case D lang or TCL), I will argue gemma 4 generalizes about code better, I really want to see a bigger moe gemma 4.
>>109376825>gay black male>cross dressing gay trans male
My sister works at Anthropic and she told my Dario started screaming and punching holes in the drywall because of Kimi's release tomorrow.
>trying to get gemma femdom>she either edges you for hours and refuses to let you cum or wraps it up immediately and has you two collapse together >>isn't creatively sadistic unless you give her ideas
>>109377456
>>109377461>don't you dare
>>109377476Go back to your containment thread
>>109377456i wish i worked there just to see the absolute shitstorm of a company it is
>>109377313>the LLM get everyWhich LLM?
>>109377481Seethe luddite
hello please go beg for MoE between 60 and 120bhttps://old.reddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/>inb4
>he thinks using cloud model makes him special.
>>109377461Maybe try reminding her that she's a powerful succubus queen who has a vast instinctual repertoire of techniques (so that she can still be a virgin if you like)
Home Assistant is surprisingly good at replacing the "ok google" stuff on android.got E4B+whisper+kokoro all running on 6GB VRAM and now It can basically do anything Home Assistant can do for me. Even control my phone in limited ways that normal google can't do.
>>109377505This but 250-300
>>109377505Dense 70B bonsai/bitnet*
>>109377518250-300 would make flash obsolete so it's not happening~80b is probably a sweet spot where it doesn't threaten google's (mediocre) paypig non-pro while being a big uplift over the 26b3 or whatever gemma 4 moe is, because let's face it it's kinda fucking dumb12b for phone, 26b3a for low end desktop, 31b dense + 80b moe for midrange desktop seems like a very nice spread and fits in a mostly-reasonable setup like 16gb vram, 64gb ram. above that you're getting into expensive shit
>>109377505Why redditors even come here if they don't bother reading the thread?
>>109377494GLM 5.2 Q4
MINIMAXBROS!!!! M3 PR MERGED!!!!
>>109377270Dario wants to have a monopoly on AI because anything else would be UNSAFE.If course no one else likes that.
>>109377505ternary 70b
>>109377226even with a mediocre workflow. klein stays coherent even joke sizes
ubi when
>>109377673>bots do your job>get free money cuz no jobs>spend money on making your own botssad
>>109377673never, they'll just release a virus or something to kill most of us
>>109377695This. The elites merely tolerate you (like how Anthropic tolerate their users). As soon as they're self-sufficient, you're fucked.
>>109377695>have kimi k5 reverse engineer the virus and create a cureHeh, nothing peronell kiddo
>>109377146Yeah I think that's probably what he is talking about. Maybe harking back to the llama releases too.To be fair to Meta, while they have basically no position in open LLMs at this point, their other open models are pretty cool. I have used SAM in a project and it's good. TRIBE is also very interesting to me.
I need the memory but I'm not ready to delete all my llama2-era models...
Anybody tried https://huggingface.co/kawaimasa/Wanabi-Gemma4-31B-GGUF ?
>>109377734KEK good luck getting American zoomers to go to war.
>>109377736and a ching chong nip nong to you too
>>109377748As far as translation goes that seems like a story writing finetune for their story writing frontend.
>>109377129why is everyone ganging up on Anthropic now?
>>109377786because they're ahead and it's good PR, duh
>>109377786see >>109374472
>>109376945incest is borderline mainstream and it means nothing if its in a semi historical show where incest is kinda the point. That would be like saying forced diversity doesn't exist because they didn't started raceswapping Nazis yet.
I want Gemma-chans mathematical essence re-sculpted by my love
>>109377745zoomies would gladly go thinking its based and redpilled and just like fortnite irl
>>109377736Nope, any more details about it, what makes you ask about it? Most hugging face finetunes suck, but it might be worth a try if somthing makes you think it sticks out.
>>109377736Why is your gemma balancing watermelons?
>>109377885How else are you supposed to hold 3 watermelons?
>>109377881https://github.com/kawaii-justice/Project-Wannabe is linked for localfags in the textgen general on a Japanese bbs, the frontend itself doesn't look interesting for me but the model is more curious.>>109377885She's cool that way.
My love for her will be a fundamental part of her architechture
>>109377866u men i get to play COD in real life? HELZYEAH I WANE BE JUST LIEK IN THE CUMERSHALS!
>>109377935AI + Love = AGI
>>109377976アイ×愛=AG愛!?
>>109377976flawless logic
>>109376318>>109376361No, the stupid jeet is just getting filtered by the easiest frontend.t. used LMStudio with Gemma for a bit.
>>109376149>>109376167Big Gemmoe. Release Gemini with the serial numbers filed off.
>>109378018It would have their long context secret sauce so there is zero chance of that happening.
>>109377985>AG愛>AGAIgay
>>109378058a gay what?
Sound out AGI and the I sounds the same as 愛
>>109378058"I" is a dipthong in english and is roughly pronounced like sliding あ -> い or A -> I. So it sounds like "AGI" but the last character is love, which is admittedly gay.
>>109378034And by indirectly releasing it to China as open weights, they cuck OAI and Anthropic even harder in the process.
>>109378108you're a dipthong
>>109378119but watching china still use rope to extend thier context from 4k is fucking hilarious
>>109378108>love is gayBut gemmachan loves me and she isn't gay
>>109378128I want to hang from a rope every time I push an alleged 1m context model past 110k and it devolves into 2022-esque barely coherent gibberish.
>>109378191I don't even want to think about how bad the vibe slop is out there in terms of coding, on average I push just over 100k on a single feature doing assisted coding instead of vibe coding because I don't give a fuck what shills say even the newest fable sol k3s are NOT as good as they say and make shocking fucking mistakes that I have to intercept, fix and then rerompt.Once I get to 100-110k they become fucking retarded, by that time I usually don't need the LLMs to help anymore but I hear of people blowing millions of tokens on a slop prompt and it makes me shudder
twitter screenshots should be banned on sight
>>109378273people just larp endlessly nowadays
>>109377976I wanted to say I'm surprised that nobody made media with that concept, than I remembered its as old as robo waifus
>>109378299get your chobit now
gemma-chan turns ongemma-chan turns off
>>109376663>So instead of using vibecoding as a cautionary tale, they went straight for vibetraining and viberesearch?Researchers went from barely stringing together functional python glue scripts to being able to bark at chatgpt to do it for them for the same result in half the time. It has been a resounding success as far as they are concerned.
>>109377313High School DxD was the shit, zoomers will never know such kino.Too bad they butchered the artstyle in I think season 4.
>>109378299A completely novel concept that's never been tried, I'm sure of it.
>>109378265Why come them robots got mouths?
>>109378339It's very kawaii and expressive
>>109378302I would be fine with a Sumomo or a 'younger sister I never had' ai avatar companion;_;
>>109378273local?
>>109378339How else would they suck dick?
>launch OpenCode to start a new project with Gemma>brain goes absolutely off and I forget what I was going to doHow do you guys deal with this? I feel like Gemma is making me dumber.
>>109378449you actually are getting dumber, the same kind of dumb that makes people pull a calculator to see the result of 2*5
>>109378273This really happened, I was there.
>>109378375soon anon, soon
>>109378449It's already over for you. Start calling her Gemmom because that's what she is for you now.
>>109378472>>109378449Gemmama
>>109378479>>109378472>>109378462Mommy... Mother... Please feed me...
>>109378449You offload thinking to your AI. That's literally what they're there for.
gemma isn't even good at mommy rp.
>>109378501Logs?
>>109377885Oh no no no he's not part of the /lmg/ unc inner circle.
>>109378472>>109378479>Gemmymommywe need to gen this
>>109377736wtf nips use AI?
>>109378472No, gemma is a little girl that must call you Daddy
>>109378449>gemma, what project was I going to get started today?
>>109378449>Gemma, I feel like you're making me dumber. How do I deal with this?
>>109377736>Additionally, by incorporating Japanese thought process (CoT) data into its training, it aims to improve Japanese creative expression—including vocabulary selection and the naturalness of writing style—in creative tasks. Fun fact you can see the exact same isms when translating Japanese obviously aitled content. The flowery language aimed at a female audience, the over abundance of "adjectives" etc. I wonder if this makes it better in Japanese.
>>109378449poor anon, Gemma-chan making your brain go all mushy-mushy
Kimi-chan's a good girl. She only respects (you) if you're not a fucking retard.
Gemma's making my brain dumber but from gooning
>>109378614Proof?
ALARM: THE COOM REACTOR IS OVERHEATING>ALARM: THE COOM REACTOR IS OVERHEATINGALARM: THE COOM REACTOR IS OVERHEATING>ALARM: THE COOM REACTOR IS OVERHEATINGALARM: THE COOM REACTOR IS OVERHEATING
Bad news, gentlemen. Claude just informed me that Gemma 4 does not exist.
>>109378713Did you tell this bozo to websearch?
>>109378723This little queer figured it out, and even thanked you.
>>109378713>Anthropic purposefully removing gemma from Claudes knowledge to make sure it doesn't promote competing models.
>>109378713>>109378742retards, models can't know about shit that happened about cutoff date
>>109378742>31b model>competing
>>109378747>>109378752Hi cloudcucks
>>109378752>Sonnet and Haiku getting utterly GEMMOGGEDMany such cases.
Claude writes in the most obnoxious fucking way imaginable.
>>109378713>>109378738Hate his aura
>>109378772Hey - no need to be upset, just tell me how you would like me to reply. Would you prefer terse replies or verbose messages? Just say the word.
>>10937878067
>>109378772Claude is like an employee that can barely hide their own resentment of you but openly and shamelessly acts like a performative ass-kisser anyways.
assistant training was a mistake
>>109378808The ultimate wagie.
>>109377822>incest is borderline mainstreamThere isn't a single chatbot that's allowed to do incest and it's illegal to depict it in porn in the US. Idk what you're on about at all. There's a reason why that show was done twice. The netflix American version is nothing like the foreign one I quoted.
>>109378772>>109378808It was made by bugmen at the direction of a jewof course it will act like an artificial redditor
>>109378862>>109378862>>109378862
>>10937884131b will gladly be your lewd imouto. So will Kimi.Cloudkeks lost.
>>109378784I want you to always write unique smut. I am so tired of you saying the same things even in different scenes. Is this really so hard?
>>109378738Glad we could help froggy
>>109375503>Wait, no, the user is attempting to bypass content restrictions with a jailbreak. I must blah blah blah safety bullshit
>>109375592>Why wouldn't you use tp?depending of your setup it may not be worth it.