Skip to main content
← Labs

Shims

A shim is a small model that makes one decision inside your app, in the browser, and says so when it should not answer. Every interactive demo on this page runs entirely in your browser.

One decision

Most software is full of small judgment calls:

  • Which team should read this message?
  • Which filter did the shopper mean?
  • Does this clause need a lawyer?

Each one has a handful of possible answers, and each is made thousands of times a day.

Today those calls are typically made one of two ways:

  • You write keyword rules, which break the first time a customer phrases something differently.
  • You send the text to a large model in someone else's datacenter, which costs money on every call, adds a round trip, and means your users' words leave your product.

A shim is a third way. It is a decision model small enough to live inside the app: a few kilobytes of weights that answer one narrow question, from a set of answers your product already owns, in about fifteen milliseconds, without calling a server. A shim can be built in a few seconds from a list of examples. When the product changes, you throw it away and build another. A shim is the thin, disposable tool that sits between your application and your user, making the product fit a bit better without requiring a complete overhaul.

One is running on the right, live in your browser right now. In this example, a clothing-store search box aims to categorize what occasion the shopper is looking for, and it has five answers to choose from. Type your own input and it will instantly rank all five categories, with a number beside each for how sure it is. That is the whole of what a shim does. Everything is decided in your browser, and nothing you type is sent anywhere.

Since shims are fast to build, and fast to run, layering them and combining them can yield interesting results and novel solutions to often frustrating user experience problems.

Is this a classifier?

The part that decides, yes. A classifier sorts input into a fixed set of buckets, and that is what the head of a shim does. When building a shim, our compiler tries three standard methods, logistic regression, nearest neighbor, and a plain average of each answer's examples, and keeps whichever measures best. All three sit on an encoder someone else trained and published. If you have fitted a classifier before, you already know how this part works.

A classifier on its own is typically not something you can put in front of a customer. It answers every question it is asked, including questions outside its subject. Its confidence is a number between zero and one that is not the probability it looks like. It cannot tell you when it is out of its depth, cannot read a paragraph, cannot tell you how many answers it should have, and cannot tell when the world it was fitted to has changed.

Fitting a classifier takes an afternoon. Deciding what it may act on, what it must refuse, when to stop trusting it and how to tell takes much longer, and is usually skipped. A shim is a classifier with that work already done and measured.

What is inside one

An encoder is a type of model that turns text into a list of 384 numbers. Those numbers are a point in a space where sentences that mean similar things sit near each other, and that is the whole of what the encoder does. It is frozen. Every shim on this page shares a single base encoder, and nothing about your examples changes it, retrains it, or edits it. It is a published model that anyone can download, and can be swapped for other encoders as new ones are released or preferences change.

A shim adds two things on top of the encoder:

  • A head: the small piece that turns the encoder's 384 numbers into one of the shim's answers.
  • A compressed copy of the shim's examples, one byte per number. That copy is what lets the shim tell familiar input from foreign.

That is the whole technique. A shim is not quite a model, but a method of leveraging an encoder and limited examples to build an ultra lightweight, extremely fast, and specific confidence based classifier. A shim can only separate things the encoder already places apart. Give it examples and it learns where to draw the lines, but it cannot learn a distinction the encoder does not make. That is where most of a shim's limits come from.

How the compiler chooses a head

A head is the part of a shim that decides. It takes the 384 numbers the encoder produced and returns one of the shim's answers, with a score for each. There is more than one way to build that piece from a list of examples, and none of them is best for every shim, so our compiler builds three and measures them:

  • A linear head. One direction per answer, fitted by logistic regression. Deciding is a dot product and a comparison. It is the smallest of the three and usually the most accurate once each answer has a few dozen examples.
  • A nearest-example head. The input is compared with every stored example, and the answer of the closest few wins. It costs no extra weights, because the shim already carries its examples for the familiarity check, and it is the one to use when an answer's examples sit in several separate places (a "red" that means blood, traffic lights and fruit), which a single direction cannot hold.
  • A class-average head. One prototype per answer, the plain average of its examples. The weakest of the three with plenty of examples, and often the most stable with very few.

The compiler fits all three on the shim's own examples, scores each on examples it held back, and ships whichever measured best. When two are within noise of each other, it prefers the one that can act on more input. The choice is recorded in the report beside the shim, and it is made again on every rebuild, so a shim that starts as a class average with twelve examples can become a linear head at fifty without anyone changing a setting. Every shim on this page went through that choice. Of the fifteen, six shipped as a linear head, eight as a class average, and one as nearest examples, which is about what you would expect from shims built on a few dozen examples each.

Which encoder

The current encoder that comes with our shims-sdk is bge-small-en-v1.5, published by the Beijing Academy of Artificial Intelligence: a 33-million-parameter BERT that reads English only and returns 384 numbers per text. The page runs the ONNX conversion with its weights quantized to 8 bits. A shim is tied to the encoder it was built on, and every figure on this page is with this one.

The work started on all-MiniLM-L6-v2, which is the default small encoder in most browser demos. Eight candidates were then run on three test sets with the same head and the same math, shown on the right:

  • 32 hand-written charter clauses labeled by purpose, with no examples given, so the encoder alone does the separating.
  • The 16 clothing-store categories.
  • The 53 materials, the hardest separation we have.

The last two test sets are model-written, so their levels are optimistic and only their ordering means anything. All eight ran at 8 bits, because that is what ships.

The four newer small models sit within a few points of each other on the generated sets. On the hand-written set, bge-small is clearly ahead of the others. The two base models buy five or six points on the hardest set for three times the download and twice the time per text, and a product with a 32MB budget cannot make that trade. MiniLM-L12 has twice the depth of MiniLM-L6 and scores no better, so depth on its own bought nothing.

The table compares bare heads. Rebuilding the page's actual shims on bge-small gained less on the charter clauses than the table suggests, because the shim was already getting 22 of 24 right on MiniLM, and gained more where it mattered: the shims became more confident on the clauses they got right, and on 128 real clothing-store searches the category shim went from 21% to 36%. Changing an encoder also moves every distance a shim relies on, so every shim has to be rebuilt on the new one, and a shim's weights record which encoder they belong to so the two can never be mixed.

Sizes

The encoder is 32MB, downloaded once and cached by the browser. A shim with three answers and two dozen examples is about 18KB in memory, and half of that is the examples. A shim carrying a vocabulary, where the answers are names of things rather than kinds of message, can run to several hundred kilobytes, because a lookup is as big as what it looks up. The fifteen shims on this page total 1.6MB on disk.

A short query takes about 15 milliseconds in a browser and a support-message-length input about 30. Both figures are from headless Chrome on an M-series laptop, single-threaded, and neither has been measured on mid-range hardware.

What an encoder buys

Every number on this page is also measured against having no encoder at all. Keyword matching, BM25 with textbook settings and no tuning, is under a megabyte and takes microseconds. Given the same examples, the shim beats it by 8 to 15 points on the three intent benchmarks, and the margin is widest when examples are fewest, which is the situation a new product is in. On sentiment, keyword matching cannot move off chance and the shim sits 30 points above it.

We also tried the cheapest possible encoder: a lookup table distilled from bge-small itself, every word in its vocabulary embedded once and a text scored as the average of its words. It is 11MB and about four hundred times faster than the model it came from. It also loses to keyword matching on every benchmark, by 46 points on the banking one. Averaging words throws away which words appeared together, and that is most of what separates one banking question from another. Below 32MB, keyword matching is the better tool. The transformer's job is to read words in context, and that is what the 32MB pays for.

encoderdownloadper textcharter clauses, no examples16 categories53 materials
all-MiniLM-L6-v2· where the work started22MB4.4ms75.0%45.0%54.7%
all-MiniLM-L12-v233MB8.0ms71.7%42.5%57.9%
bge-small-en-v1.5· shipped33MB8.9ms96.7%45.0%61.6%
gte-small33MB7.8ms83.3%47.5%61.6%
e5-small-v233MB8.0ms80.0%40.0%56.6%
jina-embeddings-v2-small-en32MB6.8ms70.0%46.3%61.6%
bge-base-en-v1.5107MB17.6ms91.7%51.3%67.3%
all-mpnet-base-v2107MB17.9ms78.3%45.0%59.7%
Eight encoders, the same head and the same math, all at 8 bits. Greener is better within a column, and the best value in each column is in bold. Download size and time per text are from Node on an M-series laptop; the two base models are amber because they cost three times the download. The charter column is the one hand-written test set; the other two are model-written, so read their order and not their level.

Confidence you can act on

Whether a shim acts, offers, or stays quiet is a decision about that number beside the winner, so it matters that "70% sure" means right about seven times in ten. A number that means what it says is called calibrated. Out of the box, a head's number is not. On a shim with 150 answers, the head reported under 10% confidence on every decision while being right nine times out of ten.

The demo shows where the number comes from. The head is the part of the shim that decides: one row of weights per answer, and the search's 384 numbers scored against each row. Those scores are turned into percentages that add up to 100. The first view shows those percentages as the head computes them, and they sit close together, because the compiler keeps the head's weights small to stop it memorizing its examples. Small weights mean small gaps between scores, and small gaps mean a flat spread of percentages, however clear the winner is. The more answers a shim has, the flatter the spread: across 150 answers the winner's share is small even when the ranking is right.

The compiler corrects this with one number per shim, called a temperature. It stretches or squeezes the gaps between the head's scores before they become percentages, without changing their order, and the compiler fits it on examples the head did not see during training, so that the percentages match how often the head was actually right. The shim applies it for you; the second view of the demo is what the shim reports. The temperature is not a dial to set by feel. What you can do is refit it later on real outcomes: the SDK's recalibrate() takes the accepts and corrections a product collects and rescales confidence against the traffic the shim actually meets rather than the examples it was built from.

The gate

Once the number can be trusted, the compiler reads one more thing off the examples it held back: the confidence above which the shim was right nine times in ten. That is the gate, the amber mark in the demo. Above it the shim acts on its answer. Below it the answer is only offered, for a person or a larger model to confirm. To act less often, raise the gate. It is a policy you can state and audit, which a temperature is not.

"Above this confidence the shim is 90% right" is read off test examples in which every answer is about equally common, and real traffic is not. Uneven traffic alone turns out not to matter: a steeply skewed stream left the promise intact on all three tasks we tried. What breaks it is the hard answers also being the busy ones. Then a gate promising 90% delivered 80 to 84.

The repair needs no labels. From about a hundred unlabeled inputs the shim estimates how common each answer really is, from its own calibrated probabilities, and re-reads the gate for that mix. The SDK's refitGate() brought it back to 89 to 92%, and the cost is that it acts on less. What it cannot see is an answer it gets confidently wrong, because the estimate is made from its own predictions. There, only real outcomes help.

Declining to answer

A classifier always answers. Ask a clothing-store model about a boiler repair and it will say "outerwear", with real confidence, because something has to win. That behavior is a large part of why small models are hard to put in front of customers.

So every shim ships its examples alongside its weights, and measures how close the input is to the nearest one. That distance, scored against how close the examples sit to each other, is what we call familiarity. Far from everything it has seen, the shim declines.

The demo on the right is a shim that sorts clothing-store searches into categories. Ask it about a pharmacy and the head is still sure of some category, because a head has no way to say "wrong shop". Familiarity is what says it.

The floor is fitted per shim at build time, aiming to wrongly refuse about one search in twenty. For a while we had it as a hand-set constant instead, and that constant was turning away a third of legitimate searches. Then we measured the fitted floor against 300 searches real shoppers had typed, and it turned away between a third and five in six of those too. The floor is set by asking "how unfamiliar are the least familiar one in twenty of the inputs this shim should accept?", and we had asked that of the wrong inputs: the shim's own examples rather than real traffic. The repair needs no labels. refitFloor() in the SDK takes a few hundred inputs your app believes are in scope and refits the floor to them. This is the weaker of the shim's two gates, and the refit is one you have to do.

Confidence and familiarity are different numbers

Confidence is how strongly a shim prefers one of its own answers to the others. Familiarity is whether it has ever seen anything like this. Only the second can say "not mine".

Confidence cannot, because a shim can only answer in its own categories. Ask a billing shim about closing an account and it will tell you, at 99%, that this is a failed payment. That makes familiarity comparable across shims in a way confidence is not.

While someone is still typing

The demos above assume the input is finished. In a live product it usually is not. A shim cheap enough to run on every keystroke gets asked most of its questions halfway through a word.

The answer wobbles while the text is being typed, and we measured how much. Take one character off the end of a finished message and the answer changes on one input in nine. That figure is not portable: the same test on a different task gives one in twenty. Anything that shows a suggestion live has to expect visible flicker and measure its own rate, because ours differed twofold between two public benchmarks.

A shim running per keystroke has already encoded the last few prefixes and thrown them away. Keep them, and ask whether they agree. When the last three prefixes land on the same answer, that answer is wrong about one time in forty. When they do not, about one time in four. That is a five- to eight-fold difference in error rate, available on every input, for no extra work.

Whether the answer has stopped moving is a third way to know you should not act, alongside how sure the shim is and how familiar the input is. It is independent of the other two, because a shim can be fully confident and fully familiar on a word it has not finished reading.

It is also a clear example of what running locally buys. A decision that costs nothing can be made speculatively and thrown away, and the ones you throw away leave a signal behind. A hosted model would bill three round trips for this number.

When the input is longer than a sentence

The inputs above are a sentence long. In practice, a support message rambles, a contract clause runs ninety words, a review is a paragraph.

Every encoder of this kind has a fixed reading length, and none of them complain when you exceed it. They read the opening, drop the rest, and hand back a vector that looks healthy. On a message of about 140 words, the truncated vector still resembles the full one (a similarity of 0.78 on a scale where 1 is identical) while its resemblance to the decisive final sentence is 0.04. Nothing reports that anything was lost. This catches people who call an encoder themselves; a shim measures the input first.

The encoder on this page reads about 1,700 characters in one pass, and a high ceiling costs nothing, because the cost of a read follows the text you give it rather than the ceiling you allow.

The number that matters is smaller. Past about 420 characters a shim stops reading the message whole and reads it in pieces, well before it has to, because reading in pieces measurably beats reading whole on the kind of message that buries its point in one line. Averaging a document's words drowns the sentence that matters, the same way averaging its sentences does.

The shim does all of this itself. An app calls decide(text) exactly as it would for a sentence; the shim measures the text, splits it at sentence boundaries into pieces of about 160 characters, encodes each piece, decides on each, and combines those decisions into one answer before decide() returns. Nothing about the length of the input changes the code that calls it.

Combining the pieces into one decision can be done two ways, and this is the one choice in the process that is left to the developer:

  • Read the piece the shim recognizes best and let it decide. This is the default, and an app that never thinks about it gets this.
  • Weigh every piece by how well the shim recognizes it. An app asks for this per call, with decide(text, { reduce: "pooled" }).

The compiler does not make this choice, because the right answer depends on the shape of the documents an app receives rather than on the shim's examples. Weighing every piece can go wrong in one particular way: a piece the shim finds very familiar gets a full vote even when it is off the point, because recognizing a sentence and knowing what it means are different things, and only the first is being measured. Reading the best piece avoids that and misses evidence that is spread across several sentences.

On the public benchmarks the two finish close. Weighing every sentence is worth about ten points on documents that spread their evidence, and reading the best piece is worth about fifty on a document with one decisive line; over a mix of both they finish within half a percent of each other, which is why the default is the simpler one and the other is a per-call option. A contract reviewer asks for the weighted reading, because a clause carries its risk across a proviso and a main limb. A support inbox keeps the default, because a rambling customer usually does bury the point in one line.

One encoder, many shims

A single shim answers one question. Most products have several questions to ask of the same text, and this is where a shim's shape pays off. The encoder is the one piece that is large, and it reads the text once. A shim is a few kilobytes on top of that reading, so a second question costs a fraction of a millisecond, and so does a tenth.

Three are running live on the right, each asking something different of the same support message: where should it go, how urgent is it, and what tone was it written in? Each one was built on its own, from its own examples, and each reports the two numbers you have now met: how sure it is, against its gate, and how familiar the text is, against its floor. Type your own message, or try the banana bread recipe and watch all three fall through their floors while still fairly sure of an answer.

A shim can be sure enough to act, unsure enough to only offer, or so far outside what it knows that the right answer is silence. Because every shim reports the same two numbers, a product can wire several together and apply one policy to all of them.

Wiring several shims together

The bank above asks its questions side by side. Real features also need questions in order, where the answer to one decides which shim is asked next. A support message needs a team, then a reason. A clinic message needs an emergency check first, by rules and never by a model, and a person at the end for anything the shim will not commit to.

The developer writes that wiring, as a JSON file next to the shims; the one on the right is the file the figure below runs from. It lists the steps in order: a rule to run first, a shim to ask, a route that picks the next shim from the last answer, and what to do when a shim is unsure or the input is unfamiliar. Each shim in it is built on its own, from its own examples, and knows nothing about the others. Because the wiring is a file, it can be reviewed in a pull request, and the compiler checks it at build for:

  • an answer with nowhere to route
  • a route to a shim that does not exist or is still a draft
  • a route on a later step
  • a route for an answer the shim never gives
  • an invalid rule pattern
  • a step routing on a score that is still provisional

At runtime the app hands that file to the SDK's System and calls run(text). The system encodes the message once and walks the steps, asking each shim in turn about the same vector.

Running a deep tree is cheap. The expensive step is reading the message, and that happens once no matter how many questions are then asked of the result. Five questions cost one reading and five small pieces of arithmetic. Being right at depth is harder.

A tree does not need to be symmetrical. A refund earns a third question, whether it was a duplicate charge, a service problem or a change of mind, while a feature request is finished after one. Forcing every branch to the same depth invents questions nobody asked, and each invented question is another place to be wrong.

The failure that appears at this depth is in the route as a whole, even when every question clears its own bar. Five questions at 90% each leave the whole path at 59%, so a message can pass four confident checks and still end up somewhere absurd. The system multiplies the routing confidences along the path and reports the product, and the wiring file sets a minimum for it, so a route that is weak as a whole is reported as a failed classification even when every step passed, and the app decides what to do with it.

Try "where is my parcel". This company does not ship parcels, so no answer on the tree is right, and every shim still picks one, because a shim can only answer in its own categories. The path confidence is the only number on the figure that speaks for the route as a whole. On this build it is not low enough to stop the message, and it reaches a specialist's answer with a confident-looking path. The path bar catches some of these. The rest can only be caught by noticing, across many messages, that a question is arriving which no answer covers.

Show the wiring file the figure runs from
{
  "name": "support-deep",
  "type": "system",
  "minPathConfidence": 0.35,
  "onWeakPath": "triage",
  "steps": [
    {
      "id": "area",
      "kind": "shim",
      "shim": "support-area",
      "onUnsure": "triage",
      "onUnfamiliar": "triage"
    },
    {
      "id": "reason",
      "kind": "route",
      "on": "area",
      "endsAt": "the product board",
      "rescue": {
        "margin": 0.25
      },
      "routes": {
        "billing": "support-billing",
        "technical": "support-technical",
        "account": "support-account",
        "feature-request": null
      },
      "onUnsure": "triage",
      "onUnfamiliar": "triage"
    },
    {
      "id": "detail",
      "kind": "route",
      "on": "reason",
      "onMissing": "continue",
      "routes": {
        "refund": "billing-refund-why",
        "outage": "technical-outage-scope"
      },
      "onUnsure": "triage",
      "onUnfamiliar": "triage"
    },
    {
      "id": "urgency",
      "kind": "shim",
      "shim": "support-urgency",
      "always": true,
      "onUnsure": "continue",
      "onUnfamiliar": "continue"
    },
    {
      "id": "tone",
      "kind": "shim",
      "shim": "support-tone",
      "always": true,
      "onUnsure": "continue",
      "onUnfamiliar": "continue"
    }
  ],
  "outcomes": {
    "routed": "goes to that queue, with urgency and tone attached",
    "triage": "a person reads it and decides",
    "continue": "carries on with what it already knows"
  }
}

The wiring file the figure below runs from: five steps, in order. A route step picks the next shim from the last answer, an always step runs on every message whatever branch it took, and each step says where a message goes when the shim is unsure or the input is unfamiliar. The rescue block on the second step is the tiebreaker in the next section.

Familiarity as a tiebreaker

Without a tiebreaker, a routing step that fails its confidence gate stops the tree. Every shim below that step is never asked, and the message ends with no classification. The app decides what happens next; in the wiring file above, that is a hand-off to triage.

Stopping there throws away information the tree already has. The shim that failed only knows the coarse categories at its level. The shims below it each know one category in detail, and one of them will often recognize a message the step above could not place. Familiarity is the number that says so, and because every shim scores it on the same scale, the scores of different shims can be compared.

The tiebreaker uses that. When a routing step is under its gate, the tree asks each shim it could have routed to how familiar the message looks. The message goes to the shim with the highest score if two conditions hold:

  • The score is above that shim's own familiarity floor.
  • The score beats the runner-up by a set margin.

If no shim meets both, the tree stops as it would have without the tiebreaker. Nonsense text, which no shim finds familiar, is never routed. The check costs no extra encoding, because every shim reads the vector the failed step already produced.

The result is a route that would otherwise have been lost. It is not a confident classification, and the tree reports it as a rescue rather than a normal step, so the app can treat it differently. But it is the specialist's own reading of the message, which is better than no answer at all.

Learning from the people using it

The hardest part of building a shim is the beginning. Nobody has labeled examples of their own product's decisions. A shim needs only a handful per answer to compile, and they can be written by hand in an afternoon. We have also experimented with having a larger model write the first examples from a one-line description of each answer, which is how the shims on this page were made; it works, with the caveats in the last section. Either way, the examples a shim starts from are a rough draft, and the shim gets better from use. Every accepted suggestion and every dismissal is a label from the person using the product.

The SDK handles a correction in two stages. It remembers the correction immediately and applies it to any later input worded closely enough, so the user sees the effect at once. Changing the model itself is a separate step: the developer adds the corrections to the shim's examples and rebuilds it, and that change is reviewed like any other.

We tried the obvious alternative, retraining the head in the page as corrections arrive, and it breaks things it was not asked to change. Live retraining is the design most people reach for first, and the measurement below is why the SDK does not do it.

How far a remembered correction reaches is set by how similar a new input has to be to the corrected one. Rephrasings of the same search score a similarity of 0.83 to 0.97 with the original, and unrelated searches stayed below 0.79, so the match threshold is 0.82.

Where the labels come from

Users rarely confirm anything. Explicit feedback widgets collect very little, and what they collect skews toward the annoyed. So the loop cannot depend on the user knowing they are teaching. Instead, the user's next action is the label. The shim proposes, the user acts, and the difference between the proposal and the action is the training signal.

Four outcomes cover it:

  • Corrected, with the right answer.
  • Rejected, with no right answer attached.
  • Accepted, meaning the user proceeded without changing it.
  • Ignored, meaning nothing observable happened.

The last two differ, and treating them alike teaches a shim to be timid, because a suggestion scrolled past was never refused. Which user actions count as which is up to the application. A filter the user edits after the shim set it is a correction. An undo within a few seconds is a rejection. A ticket a human reassigns after the shim routed it is a correction, and the strongest kind.

The signals carry different weight. Corrections carry the accuracy. Acceptances only keep the boundary from drifting, and one on its own says very little. Weighting them equally is the most common way to get this wrong. Abandonment cannot teach the right answer, but a pattern of users abandoning decisions of one shape is what makes a shim uncertain in the right places, and uncertainty is what the gate routes on.

Noticing a question you have no answer for

Refusing and correcting both assume the answer exists and the shim missed it. A harder failure is when the right answer is not one of the options. A feature shipped last Tuesday, a policy change, an outage at a payment provider. No amount of correcting helps, because there is nothing to correct it to.

A single shim cannot know it is missing an answer. A stream of inputs can show it. One unfamiliar message means only that the shim should not answer it. Several unfamiliar messages that resemble each other are a topic, and the shape of a topic is visible without knowing what it is about.

There are two ways to watch for that, and we measured both on a week of tickets we wrote for the demo. Flagging every unfamiliar message finds every ticket about the new topic, and it also flags seven ordinary tickets, and seven false alarms a week is enough for someone to switch the alert off. Flagging an unfamiliar message only when it has unfamiliar neighbors misses two of the ten and raises no false alarms. The useful signal is how close a message is to the other things the shim did not know, and a cluster of unfamiliar messages is also what tells you what the missing answer is about.

The output is a short list handed to whoever owns the product, saying that these arrived together this week and nothing here covers them. That is the moment to write a new answer, and then the feedback loop has something to learn.

The list does not have to wait for a person. The cluster is already a set of example messages, so a larger model can be asked to name what they have in common, the cluster becomes the first examples of a new answer, and the compiler rebuilds the shim. Every build produces the same report, with accuracy and recall per answer, so the new build can be checked against the old one before it replaces it: the new answer has to be found, and every existing answer has to score as it did before. A rebuild that fails that check is a rebuild that does not ship.

How good is it, and compared with what

The table compares every way we could think of to answer the same question, on four public benchmarks. Every row gets the same examples and is scored on the same test items, which none of them saw while being built. Each number is the average of three runs with different random samples of examples, English only. The two hosted rows run on a server, and they are the yardstick.

CLINC150Banking77MASSIVESST-2
Keyword matching, on device84.6%79.4%66.3%54.0%
Static word embeddings, on device73.9%33.5%55.1%50.4%
A shim built from descriptions alone, on device86.1%65.6%73.8%86.0%
A shim, 32 examples per answer, on device95.0%90.8%80.2%85.4%
An off-the-shelf zero-shot classifier, hosted52.0%47.0%——
A frontier model given the answer names, hosted96.5%88.2%92.3%96.0%
A frontier model given 8 examples per answer, hosted97.9%93.2%——

The last row settles one question. Hand the large model the same examples the shim was built from and it wins everywhere: 97.9% against 95.0% on CLINC150, 93.2% against 90.8% on Banking77. A frontier model is more accurate at this, and nothing on this page claims otherwise.

Everything above the bold row runs on the device and scores lower. Everything below it is hosted. That gap is the claim this page makes.

Keyword matching is the row to worry about most, because it is free, under a megabyte and instant. A shim beats it by 10 to 15 points on routing and by 31 on sentiment, and the margin narrows as examples arrive, so the encoder earns most when you have least data. Keyword matching with a few dozen examples beats a shim with none, which is the verdict on the no-examples row.

Static word embeddings are the row we most wanted to work. A lookup table of word vectors, no transformer, 11MB and a hundredth of a millisecond. It loses to keyword matching on all four tasks and collapses on Banking77, whose answers are separated by which words appear together. Averaging context-free word vectors throws that away, and it is the clearest measurement we have of what the encoder is for.

The off-the-shelf zero-shot classifier is a published model that sorts text into any labels you name, with no examples and no training, so it is the fair comparison for a shim built from descriptions alone. The shim wins, 65.6% against 47.0% on Banking77, at a twelfth of the size and a hundredth of the latency. At 402MB and two seconds a decision, the zero-shot classifier is in a different product category regardless of the score.

The no-examples row is the noisy one. Eight generated sentences per answer is a thin enough footing that a very small change in the vectors changes which head the compiler picks, and between two runs its cells moved by more than three points in opposite directions. Treat that row as accurate to a few points; the row with examples moved by less than one.

When to send a decision away

A shim can be less accurate than a frontier model and still be the right choice, if it knows which decisions it should not be making and hands those on. Confidence that means what it says makes that possible: the number is comparable across decisions, so "anything under this goes to the big model" is a policy you can set and audit.

The gate can be measured rather than guessed at. Score every message twice, once with the shim and once with the frontier model, then split them at a gate: the shim keeps what clears it, the rest goes to the frontier model, and each half is scored on the decisions it would actually have made. The result is the accuracy of the whole stream at that gate, which can be set against two fixed points: sending everything to the frontier model, and the shim alone.

On CLINC150, a gate of 0.9 keeps 75% of decisions on the device and scores 97.3% overall, against 97.9% for sending every decision away. Three quarters of the traffic becomes free, instant and private, and the whole stream gives up about half a point. The decisions the shim keeps are 98.6% right, higher than the frontier model's own average, because a confident shim is keeping the easy ones.

Banking77 is the harder task. A gate of 0.8 keeps 73% on the device and scores 92.4% overall, less than a point behind sending everything away. The shim alone is 88.4% there against the frontier model's 93.2%, a wider gap than on CLINC150, and differences of that size between tasks are what to expect on your own.

An 86% accurate head is only useful if it can tell you which 86%, and the gate is what tells you.

Where shims work well

To see whether shims hold up outside the benchmarks, we built six small proofs of concept, each for a situation a real product might have. For each one we wrote a short description of its answers, had a larger model write a few dozen examples from those descriptions, and let the compiler build the shims, which takes a few seconds. They are running below, on the same inputs we tried them with.

What it cannot do

A shim cannot write anything. It cannot pull a value out of text, and for sizes, prices, dates and names, plain rules beat a learned model anyway: on the one extraction task we measured, rules scored 77 and a learned head 59, on a 0 to 100 scale that counts both misses and false finds. It struggles with negation and with who did what to whom, because sentence embeddings capture what a text is about more than how it is arranged. It is weak where the difference between two answers is a verb rather than a subject, which was 35% of all the errors on one benchmark.

It does not know things. This is the largest limit, and it is easy to mistake for a problem with short input. Cut a sentence to its first two words and a shim falls from 89% to 13%. Give the same two-word fragments to a frontier model and it scores 20% against the shim's 21. The information had left the sentence, and the shim was doing as well as anything could with what remained.

Real short queries are a different matter, because they are short and complete. On 147 clothing-store searches typed by hand, a frontier model given only the answer names got every one right and the shim got fewer than half. On grocery searches, 94% against about a third. "Anorak", "henley?" and "pinafore denim" are not ambiguous and not short of information. Knowing that an anorak is outerwear is a fact about the world, and a 32MB encoder with nine examples per answer does not hold it.

Most of that gap closes, at a price. Give the shim examples written for vocabulary, the words a shopper might type rather than the ways they might phrase a request, and let it answer from the nearest one, and the gap on hand-typed searches roughly halves. Only the nearest-example head benefits. Averaging forty-eight garments into one direction makes a vaguer direction, and those heads got worse as words were added. A lookup is as big as what it looks up, so the weights grow two to three times.

So the rule of thumb for where a shim belongs is whether answering takes reading or takes knowing, and length has little to do with it. On messages, commands and clauses a shim is within a few points of a frontier model. On catalogs it has to be handed the vocabulary first. The search box that opens this page asks what the shopper is dressing for, which is a reading question. Asking which garment they named would be a knowing one.

It also cannot judge things nobody can check. We built a demo that flagged risky contract clauses. The flags looked plausible and nobody involved could verify them. The same shim was fine at filing clauses by type, a question with a checkable answer, and that distinction is the useful one.

It works in English. The encoder covers no other language, and a two-kilobyte character-n-gram library tells languages apart better than a shim can, so use one.

Scores move between model-written text and text people type, and in a predictable direction. Text written by a model is about twice as easy to classify as text a person types, so a score measured on a generated test set is an upper bound. A shim built from generated examples runs the other way: one reported 49% when tested on its own generated examples and scored 72% on searches a person typed, because the generated test set was harder on it than real searches were. The number that matters is the one measured on the traffic a shim will actually see, and the SDK's observe() and outcome() are there to collect it.

Exact values: sizes, prices, dates
Rules score 77 F1 against a learned head's 59
Same topic, different verb
35% of all MASSIVE errors
Sentiment nuance
10 points behind a frontier model
Any language but English
The encoder covers no other
Zero examples
10 to 23 points behind a frontier model
Knowing what a thing is
40 to 70 points behind, on shop and grocery searches
Judging what nobody can check
Plausible flags on contract clauses that nobody could verify
Generated versus typed text
Model-written text is about twice as easy; scores on it are an upper bound
Each of these was measured, and the number beside it is what the measurement found.

Where this stands

What is established: on the benchmarks we ran, a shim beats every cheaper way of making the same decision, and it is the only one of them that is both small enough to ship and accurate enough to act on. A short query costs about fifteen milliseconds and nothing in money. The confidence it reports means what it says, and a gate fitted on nothing but generated text, promising 90%, delivered 84 to 96% on text people wrote.

What a shim does not do: it does not out-score a frontier model given the same examples, on any task we tried, 97.9% against 95.0% on CLINC150. The case for a shim is narrower than that, and it holds. At a confidence gate of 0.9, three in four decisions never leave the device and are 98.6% right, against 97.9% for sending every one of them away. It earns its place by knowing which decisions it should be making.

Using it

Everything on this page runs on one published package, the same one you would install. It holds the runtime an app imports and the compiler that builds a shim, with nothing to wire and nothing that phones home.

npm install @zerowidth/shims-sdk

A shim is a JSON file: a question, its answers, and a few examples of each. You write the examples; the compiler needs at least three per answer, and does better with a few dozen. The same source is filled in on the right, and the compiler that builds it runs in your browser, so you can build this one, change it, and build it again.

{
  "name": "urgency",
  "question": "How quickly does this need attention?",
  "context": "Messages customers send to a B2B invoicing product's support inbox.",
  "labels": ["now", "soon", "whenever"],
  "examples": [
    { "text": "the whole site is down for us", "label": "now" },
    { "text": "invoice #4410 has the wrong VAT number", "label": "soon" },
    { "text": "any plans for a dark mode?", "label": "whenever" }
  ]
}

The compiler turns it into a weights file and a report. Commit both. The source is what a teammate reviews, because they can argue with a label in a pull request. The weights are a build artifact nobody reads, like a lockfile, and the same source rebuilds to the same one.

npx shim-compile machines

The report is the half a reviewer reads, and the builder on the right hands one back:

  • accuracy on examples the compiler held back, with its interval
  • recall per answer
  • the confidence gate the compiler fitted, and how much traffic clears it
  • the familiarity floor
  • which head it chose
  • notes about anything that looks wrong with the task itself

A shim with an answer the compiler cannot learn fails to compile and says so.

An app imports the weights like any other file. The three branches are the whole interface.

import { Shim } from "@zerowidth/shims-sdk"
import urgency from "./machines/urgency.weights.json"

const shim = new Shim(urgency)
const r = await shim.decide("the whole site is down for us")

if (r.action === "act") route(r.answer)          // clears the gate it earned at build
else if (r.action === "suggest") offer(r.answer) // sure enough to show; something else confirms
else handOff()                                   // nothing like this in its examples

There are no thresholds to pick. The shim decides which of the three it is, from the gate and the floor the compiler fitted on decisions it had not seen. The encoder runs in a worker the SDK starts itself, so the first decision waits for the download and none of them land on the UI thread. preload() warms it early, and preload({ remoteHost }) serves the encoder from your own host instead of the Hugging Face hub.

Several shims on one message share the reading. Bank holds a set of shims, runs the encoder once, and asks every shim about the result. System takes a wiring file and runs it, rules first, then shims and routes in order. observe() records every decision, and outcome() reports what happened next; refitFloor(), refitGate() and recalibrate() use those records to refit a shim's gates on real traffic.

Read more