20M params · runs at the edge

Small models.
Real speech.
No cloud.

Squeal Studio builds Russian-first language models small enough to run on the device already in your pocket — no API key, no data center, no bill that scales with usage.

20Mparameters
1,024token context
Apache 2.0license
why small

The industry keeps scaling up. We're scaling down.

01 — DESIGN

Most labs answer every problem with more parameters. We think that solves "can it do everything" while ignoring "can it run where you need it."

02 — LANGUAGE

squeal_ai_20m uses a 24,000-token vocabulary built around Cyrillic — Russian first, the language most "multilingual" models treat as a rounding error.

03 — DEPLOYMENT

Grouped-query attention keeps inference cheap enough for phones, Raspberry Pis, and browser tabs — not just racks in a data center.

the lineup

squeal_ai_20m

base

pretrained

Raw pretraining on Russian web text, education-focused corpora, and Wikipedia — no instruction tuning.

ArchitectureQwen2.5-style, GQA
Hidden size352
Layers / heads8 / 8 (4 kv)
Eval perplexity57.85
View on Hugging Face ↗

instruct

chat-tuned

Fine-tuned from base to follow instructions and hold a conversation.

Base modelsqueal_ai_20m-base
Context1,024 tokens
Runs onCPU, edge, browser
LicenseApache 2.0
View on Hugging Face ↗
zero install

Try it in your browser.

Runs entirely client-side through ONNX Runtime Web. No server call, no data leaving your machine — the model downloads once and runs from cache after that.

squeal_ai_20m-instruct.onnx cpu · wasm
idle — press run to load the model
~20M params · int8 · ~20MB