NEUON AI SDN BHD NEUSLM — small language models

NEUSLM

Three use cases, one engine

Context Mining · OCR+SLM · Self-hosted SLM

  1. 01 Context Mining
  2. 02 OCR + SLM
  3. 03 Self-hosted SLM

The slide

Three things a small model does with a document.

NEUSLM is the engine under the rest of the constellation: small, domain-tuned models that read what an institution already holds and write it into records the institution can query. This slide deck contain — three pictures of that engine at work, on documents that came off a counter, out of a herbarium and off a laptop in the office.

None of the three is a demonstration built for a slide. They are a supermarket receipt with its card slip, a Sarawak herbarium voucher, and a locally hosted model answering a question about the company that hosts it.

NEUSLM — Context Mining, OCR+SLM, Self-hosted SLM
NEUSLM — Context Mining, OCR+SLM, Self-hosted SLM

The use cases

Three documents. Three different jobs.

Each one answers a different question, and they build: read a document, read a record, then ask the whole pile something. Press a card to read it.

Under all three

From document to database.

  1. Reconcile

    Context Mining

    A photograph of each document → field extraction on all of them → one cross-checked record.

    → A reconciled record

  2. Read

    OCR + SLM

    Photograph of the sheet → text detection → field-by-field extraction → typed values.

    → A database record

  3. Ask

    Self-hosted SLM

    The institution's own documents → a local index → a small model on local hardware.

    → An answer you can check

The same four steps, three times

Ingest

Scans, forms, photographs and unstructured content enter from the partner's existing workflow.

Understand

A domain-tuned small model reads the content in its own professional context.

Structure

Findings are written into structured database records rather than left as free text.

Improve

Feedback loops retrain on partner-owned data — improvements accrue to the partner's sovereign asset.

What the three share

Small, and tuned to your documents

Not a general model asked politely about a receipt. Models tuned to the partner's own vocabulary, forms and layouts — which is what makes the difference between reading a document and reading THIS document.

It runs where the documents are

On-premise, at the edge, or in-jurisdiction. Quantised open-weight models are the reason that is affordable rather than aspirational: the hardware is bought once and the questions after that are free.

Structured out, and checkable

Records rather than paragraphs, and sources rather than assertions. Everything above is only useful if the next person can query it and the one after that can audit it.

On-premise, at the edge, or in-jurisdiction — see Sovereign Applied AI for what that commitment means, and NEUSLM for the chapter it belongs to.

Next

Bring your own documents.

The shortest way to know whether this works on your records is to run it on your records. Send a handful of the forms you actually receive — the messy ones, not the clean ones — and we will show you what comes out the other side.

  1. A sample. Twenty or thirty real documents of the kind you have most of.
  2. A field list. What you would want to be able to query, in your own words.
  3. A read-out. What was extracted, what was not, and what it would take to close the gap.