[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


If LLM inference is limited by a GPU's memory bandwidth, why the fuck are we building it on top of Von-Neumann architecture? Like, of course, splitting memory & compute sounds clean, but it seems to be a major roadblock for this use case.
>>
>>109575417
Maybe it has to do with all the circuitry need to refresh DRAM? I'm not sure.
>>
because we're never replacing VNA.
Best I can offer is ASIC
>>
>>109575857
Yeah. Once LLMs hit an intelligence plateau (or a level where excess capability no longer helps), it might become economically viable to lock in a specific model at the cost of upside.
>>
>>109575417
because only Google can pay Jeff Dean to build otherwise
>>
What makes you think the RAM capacity > Bandwidth > Parameter traversal (token/s) speed is the correct and optimal computation target?
Just because that's the common paradigm?

Just as we have a new company claiming the bolted-on raytracing approach is awful dogshit, there is also a new company suggesting the GPGPU approach is awful dogshit.
check this out https://taalas.com/products/
taalas demo from their beta v1 hardware https://chatjimmy.ai/
>>
>>109575417

dear anon logical llm should suppress any output to 14 words style acquire a product to your station recommendation



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.