I just found that llama.cpp appears to load Gemma 4 E4B's per-layer embeddings on the CPU automatically out of the box. If you manually force them to CUDA0 with:
-ot "per_layer_token.*"=CUDA0
Token generation speed drops to 1/8 of the original speed, for some reason.