Gemma 4 Nano — TinyStories Generator

A 37M parameter language model built from scratch using the Gemma 4 architecture (dual head dims, shared KV cache, K=V attention, proportional RoPE, logit softcapping, QK norm, zero-centered RMSNorm).

Trained on the TinyStories dataset with a custom 8K-vocabulary SentencePiece tokenizer.

Enter a prompt below and click Generate to create a children's story!

50 500
0.1 2
0 200
Examples

Model: lakhera2023/gemma4-nano-tinystories | Architecture: Gemma 4 (37M params, 20 layers, 8K vocab) | Training: Custom SentencePiece tokenizer on TinyStories