Gemma 4 Nano — TinyStories Generator
A 37M parameter language model built from scratch using the
Gemma 4 architecture (dual head dims, shared KV cache, K=V attention,
proportional RoPE, logit softcapping, QK norm, zero-centered RMSNorm).
Trained on the TinyStories
dataset with a custom 8K-vocabulary SentencePiece tokenizer.
Enter a prompt below and click Generate to create a children's story!