• 119 KB model
  • no GPU, no numpy
  • runs in AWS Lambda

Small model. Fast second opinion.

KinModel (the code calls it KinShield-Lite) is our own tiny scam-text classifier. It turns words into TF-IDF features, scores them with logistic regression, and fits in 119 KB. Inside KinBot it sits next to the cited Bedrock check as a labelled second opinion.

Experimental · trained by us on synthetic data · not a fine-tuned LLM

Live on AWS

Score a message, two ways

Try an example:
0 / 2,000

Your text is sent once to the KinBot check on AWS Lambda and then forgotten. Nothing is logged or stored.

How it works

From words to a score in four steps

  1. 1 · Features

    Word 1–2-grams and character 3–5-grams, weighted with TF-IDF. Character grams catch spellings like “g1ft card”.

  2. 2 · Classifier

    Logistic regression with balanced classes, calibrated with a sigmoid so the output reads as a probability.

  3. 3 · Squeeze

    Weights quantised to int8 and exported as JSON: 119 KB, against 215 KB for the float model.

  4. 4 · Serve

    A pure-Python scorer (lite.py) inside the KinBot Lambda. No numpy, no GPU. It matches the scikit-learn int8 model within 1e-9 on 52 test texts.

Results

The numbers, and why to doubt them

Positive class is "scam". Copied from our generated RESULTS.md, seed 0. Please read the caveats before quoting any of it.

Experimentn testAccuracyPrecisionRecallF1
A. Train on 600 synthetic → test on 47 KinShield scenario scripts470.9790.9551.0000.977
B. 5-fold cross-validation on the 47 scenario scripts471.0001.0001.0001.000
C. Synthetic 80/20 hold-out (easy, same templates)1201.0001.0001.0001.000
  • Experiment A got one benign script wrong (a false alarm) and missed no scams. With only 47 scripts, one example moves accuracy by about 2 points.
  • The test scripts and the synthetic training phrases were written by the same team, so A is likely optimistic. Treat it as a smoke test, not proof it works on real scams.
  • C is template data and is expected to look perfect. It says nothing about real messages.
  • No real scam transcripts, real victims' data or third-party test sets were used.

Runtime: the float scikit-learn model took a median 1.27 ms per text on a laptop CPU. int8 and float agree on every scenario label (largest probability gap 0.10). The Lambda scorer's latency hasn't been measured separately.

Research

We fine-tuned gpt-oss-20b, then shrank it again

Separate from KinModel above. Using an NVIDIA DGX Spark, we fine-tuned gpt-oss-20b (QLoRA, 90 minutes) on 1,201 synthetic calls labelled by gpt-oss-120b. We kept only the labels that matched what each call was written to be. The fine-tuned model then taught a 22M-parameter classifier, KinShield-Tiny v3. The live app still uses the stock model.

Model, on 63 hand-written test calls (31 scams, 32 safe)CorrectScams rated HighSafe calls flagged
gpt-oss-20b, stock (what the app uses)61/6329/310/32
gpt-oss-120b, stock (about 6× larger)62/6330/310/32
gpt-oss-20b, fine-tuned by us62/6330/310/32
KinShield-Tiny v2, taught by the stock 20b57/6325/310/32
KinShield-Tiny v3, taught by our fine-tune60/6328/310/32
  • The gains are small: one test call for the fine-tune and three for Tiny. With 63 calls, one call moves accuracy by about 1.6 points. Read it as "matches the 6× larger model", not "beats it".
  • All training calls were written and labelled by AI models, and the test calls were written by our team. None are real calls.
  • The fine-tuned 20b takes about 5 s per call on the DGX Spark and isn't hosted publicly. Tiny v3 runs on its own AWS Lambda in about 40–80 ms, but gives scores, not quotes.

Download: kinshield-20b and kinshield-tiny-v3 on Hugging Face.

Honest limits

Why KinModel is the second opinion, not the first

Synthetic training data

600 template-generated examples. It will miss scams that don't look like our templates.

Keyword-driven

A real bank alert full of scary words can score high. That's why KinBot always shows the cited check first.

English text only

No audio, no other languages, and no sender details like the phone number or link.

Not an LLM

It's a small classifier we trained, not a fine-tuned language model. It gives a number, never a reason.

What's next

Where KinModel is going

  1. NowSecond opinion in KinBot

    Shown next to every cited check, clearly labelled experimental.

  2. NextReal, consented examples

    Retrain and re-test on messages people choose to share, with a held-out set we didn't write.

  3. ThenOn the phone

    Run the 119 KB model on-device: faster, more private, and no cloud cost per message.

  4. LaterPre-filter for KinVoice

    Score call turns cheaply, and only call Bedrock when the model is unsure.

Want the full answer
with quotes?

KinBot shows KinModel's score next to the cited check, then helps with what to do next.