Own language model · runs in your network

An AI assistant you can trust with internal documents.

Kalorad drafts replies to customer requests, finds answers in your regulations and writes letters and code. It runs on your server: requests and documents never go to foreign AI services.

Ukrainian · English · Russian · code — OpenAI-compatible API

Kalorad · Chaton-prem
you
Що таке суверенна модель? Три речення.
KaloradLM 683M
Message Kalorad…↵
Processes · mi300x-01R→U→T→O→R
pretrain-1b24 GPU 0
96 %
model-server :8000
41 req/min
embedder memory index
312 facts/s
cycle-manager stage T · promote in
06:12:44
KaloradLM · 683M · neural activitylayer 17 / 24
synapses 1 920firing 1 240 / sheads 24 · kv 4
context 4 096precision bf16 · fp32 mastersexternal calls 0
Generating · next-token distributionT = 0.7
pretrain.log · MI300Xlive
Memory · sqlite-vec4 ms
recallWhat did we decide about data residency?
every message · every fact · local · yours
683Mparameters 13.1Btokens 43 hone MI300X $86compute source: pretrain.log · 16 750 steps × 786 432 tokens
What it does

Routine text work, done in seconds.

For the pilot you pick one scenario — the one where your people lose the most time.

Replies to customer requests
Drafts an answer by your rules and in your tone. The operator checks it and sends — instead of typing from scratch.
Search across regulations
Answers questions about internal rules and instructions and shows where the answer comes from.
Letters and documents
Drafts letters, memos and standard documents from a few lines of input.
Sorting requests
Gives each incoming request a topic and a confidence score. Unsure cases go to a person, not to the wrong queue.
Code and scripts
Writes and explains code, SQL and automation scripts. Code is almost half of the training data.
Your systems stay as they are
The server speaks the OpenAI API, so existing tools and integrations connect without rewriting.
Under the hood

Written from scratch. Every layer is ours.

No forked checkpoint, no borrowed tokenizer. The architecture below is implemented in PyTorch in this repository, tested on CPU and verified numerically on ROCm.

Architecture
Dense decoder
hidden 1536 · 24 layers · GQA 24 : 4 heads · SwiGLU 4096
Positions
RoPE + YaRN
context extension without retraining; RMSNorm pre-norm
Optimizer
Muon + AdamW
Newton–Schulz orthogonalised updates for matrices, AdamW for embeddings and head
Schedule
WSD + curriculum
warmup – stable – decay; data mix shifts by phase
Corpus · 12.5B tokens · own SentencePiece tokenizer
en 35 · uk 25 · ru 22 · code 18
EnglishUkrainianRussianCode
Precision
fp32 masters
bf16 autocast for compute; masters stay fp32 so small updates are not lost
Verification
3.3e-06
CPU vs GPU numerical parity on fixed input; 9 training bugs found on hardware, each now a test
Live inside Kalorad

Processes, memory, learning — visible.

An assistant that remembers everything and learns from everything needs an operator’s view. This is what the node is doing right now.

Processes
Memory
Training
node mi300x-01uptime 00:00:00cycle R→U→T→O→R
pidprocessstategpu / cpurate
4112pretrain-1b24kalorad_model/pretrain.py · step 16 750training
88 041 tok/s
4188model-serverOpenAI-compatible · :8000serving
41 req/min
3970corpus-builder14 workers · SentencePiecetokenizing
1.2M tok/s
4201embeddernomic-embed-text · memory indexindexing
312 facts/s
4230memory-compactorsqlite-vec · nightlyscheduled
02:14:09
4007cycle-managerR→U→T→O→R · promotion checkstage T
06:12:44

GPU 0 · utilisation

Self-improvement cycle

Rrun
Uupdate
Ttrain
Ooptimise
Rretire
VRAM178 / 192 GBpower612 Wcheckpointstep_0016750loss4.8819
recallWhat did we decide about data residency?
index sqlite-vec · embeddings nomic-embed-text · local
latency 4 ms · corpus every message, every fact · retention yours

First run · what the curve says

Steps 0 – 13 750: loss flat at 6.1. The dataset shifted labels once more than the model did — for 37 hours it learned to predict the token after next. Found in the curve, fixed on the node, now guarded by test_training_objective.py.
Steps 13 750 – 16 750: 6.1 → 4.88. Same weights, correct objective. Perplexity 132. The next run starts with the fix, the reference Muon and fp32 masters.
tokens13.1Bthroughput88 041 tok/swall43 hcost$86
Console preview. Training figures are from the run log; process rates are illustrative.kalorad_model/reports/run_700m_20260910
Security

Your data stays in your building.

Not a wrapper around someone else’s API. The model, its training and the answer server are ours — so nothing has to leave your network.

CLIENT PERIMETER · CLOSED NETWORK Employees Documents · DB KaloradLMweights · memory · loop Vendor API0 calls
01

Works without internet

Kalorad is installed on your server in a closed network. Requests and documents are processed right there.

02

No foreign vendor

The model is trained from scratch by our team. There is no third-party licence that can be revoked or changed.

03

Your data is not our dataset

Tuning for your tasks happens on your server and stays yours. We do not train the general model on your data; updates reach you as packages.

04

Encryption and an NDA

Data is stored encrypted, only your staff has access, and we are ready to sign a non-disclosure agreement before the pilot.

Free pilot

Eight weeks. No obligations.

The pilot is free and creates no financial obligations. It starts only if the answers suit you on your own examples.

Before the pilot

Check on your examples

You send 20–30 typical requests, we show the answers. Nothing is installed yet.

Week 1

Scenario and success criteria

We agree on one scenario, how success is measured and who takes part.

Weeks 2–3

Installation on your server

Deployment in your closed network and tuning on your documents.

Weeks 4–7

Real work

A group of your employees uses Kalorad every day and rates the answers.

Week 8

Report and decision

Results against the agreed criteria and recommendations. What happens next is your call.

What we need from you

  • One contact person on your side
  • A Linux server in your network with a 16 GB+ GPU (for the pilot a CPU server also works, just slower)
  • Documents for the chosen scenario
  • 5–10 employees to test it
Pricing

Start free, decide on results.

Start here

Pilot

Free

8 weeks · one scenario

  • Check on your examples before the start
  • Installation and tuning by us
  • Report with measured results
Discuss a pilot

On your server

By agreement

Annual licence after the pilot

  • Model updates as packages
  • Tuning for new tasks
  • Special terms for the first partners
Write to us

Personal use

Soon

Web chat and apps for Mac, Windows and iPhone

  • Opens together with the production version
  • Free plan to start
  • Your ratings improve the model — only with your consent
Write to us
Where we are

What is ready today.

683M
parameters in the first fully trained model (September 2026)
13.1B
tokens of training data behind it
1.24B
parameters in the production version — expected in November–December 2026

Pilots run on the production version, after the check on your examples.

Questions

Frequently asked

Does Kalorad need internet access?

No. After installation it works entirely inside your network. Updates arrive as packages that your team installs.

Which languages does it understand?

Ukrainian, English and Russian, plus program code. All of them were in the training data from the start.

How does it connect to our systems?

Through an OpenAI-compatible API. Most tools that work with such APIs connect by changing the server address.

What hardware is needed?

A Linux server with a GPU with 16 GB of memory or more. For the pilot a CPU server is possible too — answers are just slower.

What happens to our documents?

The pilot runs on your server, so your documents stay with you. For the check before the pilot you send us 20–30 examples — pick ones without personal data.

How much does it cost after the pilot?

An annual licence for your server. The price depends on the scope and is agreed with you before the pilot ends; the first partners get special terms.

Can we try it before a pilot?

Yes — that is the check on your examples: you send typical requests, we return the answers. Nothing needs to be installed on your side.

Contact

Tell us which task takes the most time.

We will propose a pilot scenario and a list of examples for the check. Write in Ukrainian, English or Russian.