If LLM inference is limited by a GPU's memory bandwidth, why the fuck are we building it on top of Von-Neumann architecture? Like, of course, splitting memory & compute sounds clean, but it seems to be a major roadblock for this use case.
>>109575417Maybe it has to do with all the circuitry need to refresh DRAM? I'm not sure.
because we're never replacing VNA.Best I can offer is ASIC
>>109575857Yeah. Once LLMs hit an intelligence plateau (or a level where excess capability no longer helps), it might become economically viable to lock in a specific model at the cost of upside.
>>109575417because only Google can pay Jeff Dean to build otherwise
What makes you think the RAM capacity > Bandwidth > Parameter traversal (token/s) speed is the correct and optimal computation target?Just because that's the common paradigm?Just as we have a new company claiming the bolted-on raytracing approach is awful dogshit, there is also a new company suggesting the GPGPU approach is awful dogshit.check this out https://taalas.com/products/taalas demo from their beta v1 hardware https://chatjimmy.ai/
>>109575417dear anon logical llm should suppress any output to 14 words style acquire a product to your station recommendation