>>109805812
Those numbers are without --split-mode explicitly set. All that I ran for both backends was:
./llama-cli --model /path/to/model.gguf
Now that you bring it up, I need to try with the flag set either way for both backends, and my old flags when I had the M10s in the system (modified to remove CUDA specific shit, of course).