A 0.6B model decodes a 31B model's hidden state
gemma-4-31B was given a photograph. At layer 47 of 60, during the prefill and before a single token had been generated, one 5376-dimensional vector was taken, pooled over the image token positions. Qwen3-0.6B, a text-only model that has never seen a photograph, receives that vector and nothing else.
gemma's side is precomputed at full precision, so nothing here waits for a GPU and you can take the state apart as much as you like.