V

Vashtra‑0.6B

A small language model that actually knows machine learning
Runs 100% in your browser · no server Model ↗ Data ↗
Checking your browser for WebGPU…
Explain why batch normalization helps training.
LoRA vs full fine-tuning: when does each win?
When should I use focal loss instead of cross-entropy?
What causes exploding gradients in RNNs?