In short
We built our own AI decision model, J33-Jev, on Google's free EmbeddingGemma 2 and trained it on 1,080 practice cases. On 2,000 test decisions it saved 22.5 pieces of work per 100, the best result we have measured on this test. ConvAI Innovations' Laya was more accurate and made far fewer mistakes. Ours saved more because we fitted its confidence on held-out practice cases. How sure it said it was then matched how often it was right, so it could safely decide alone more often. Both it and Laya had studied practice cases from this same test, which TypeSafe's Jev, used as delivered, had not. Building your own AI decision model is worth it when you have a record of past decisions, with the answers your staff settled on, to train it on. Test an off-the-shelf tool on your own past cases first.
What we tested, at a glance
| Model | Made by | Size | Licence | Needs a GPU? | Good straight out of the box? | How close to real work? |
|---|---|---|---|---|---|---|
| J33-Jev (our model, simplest version) | J33, built on Google DeepMind's EmbeddingGemma 2 | 271.6 million | Apache 2.0; free to download from Hugging Face | No. Same answers on a laptop processor, just under a second per case | There is no out-of-the-box version: we built it by training on this test's practice cases | Best result (+22.5), after studying this test's practice cases. Promising; train and test it on your own cases |
| Laya, trained by its makers on practice cases | ConvAI Innovations | 421 million | Apache 2.0 | No, ran on an ordinary server (last article) | The general version, no (36.1% right). This version had studied the same practice cases as ours | Strong (+17.1) after re-tuning how sure it says it is, as its documentation advises. The most accurate, and the fewest mistakes |
| Jev | TypeSafe | Not published | Online service, not downloadable | No, it runs on TypeSafe's servers | Yes, used exactly as delivered | Strong (+13.0) with no training on this test's practice cases. Your data has to leave your company |
| Julia-1 | Supersonic Labs | 144 million | Apache 2.0 | No, ran on an ordinary server (last article) | Accurate (73.0%), but its confidence scores were not a good guide to when it was right | Not for deciding alone with its scores as delivered (−43.5) |
Size is how many learned values the model has. The number in brackets is the work each tool saved per 100 decisions, explained below. We also built four other versions of J33-Jev; they are in the results table.
Think of it as training your own new hire
Buying an off-the-shelf model is like hiring someone experienced from another company. Building your own is like taking a bright graduate and training them on your own past cases. Either way, the skill that matters is knowing when they might be wrong, and asking. The question for a business is whether the training pays off.
What we did
We used the same public case set as last time, typed-decisions. It covers four kinds of work: customer service, invoices, security alerts and checking on AI assistants. Each case, such as one support ticket, comes with five questions. The test has 400 cases, which makes 2,000 decisions.
The cases are made up for testing. The "right answers" were set by another AI model, not by people, so "right" means "agreed with that model". That model agrees with itself 73.5% of the time. A score well above that partly means a tool has learned the labeller's habits.
We kept the same score, agreed before looking at any results.
- The AI decides alone and gets it right: +1, a piece of work a person no longer does.
- The AI decides alone and gets it wrong: −5, because someone has to find and fix the mistake.
- The AI hands the case to a person: 0, the same as today.
One mistake cancels out five right answers, so the AI should only decide alone when it is at least 83.3% sure.
We built J33-Jev in two parts. The base is EmbeddingGemma 2, a free model from Google DeepMind that turns text into numbers a computer can compare. On top, we added a small scoring layer that reads each case and question and scores every possible answer. Some versions also put an extra layer between the base and the scoring layer. We followed the design of Laya, whose makers publish their code. The name J33-Jev refers to the Jev-style format: typed questions in, a confidence for each answer out. J33-Jev is not affiliated with or endorsed by TypeSafe, the maker of Jev.
We fine-tuned the whole model on the test's practice cases. Every part of it learned from them, EmbeddingGemma 2 included, not just the scoring layer. The test comes with 1,200 practice cases, kept separate from the 400 test cases. For the four text-only versions, we trained on 1,080 of them and kept 120 aside. We used those 120 to fit the model's confidence setting (this is called calibration), so its "how sure I am" matches how often it is actually right. We never tuned anything on the test cases.
We tried five versions, to see what made a difference:
- First version, with an extra layer.
- First version, refitted: the same model and answers, with only its confidence setting fitted again, this time to the single best answer for each case. The first fit had left it underselling itself.
- No extra layer: the simplest version, with the scoring layer reading the base model directly.
- Bigger extra layer: a larger version of that extra layer.
- Also trained on images: trained on the business cases plus questions about photos from the public COCO photo collection.
Every version had one training run. Small gaps between versions, of a few percentage points right or a few pieces of work per 100, could change if we ran them again.
What we found
One unit is one decision a person no longer has to make. So +22.5 means that for every 100 decisions, the tool took 22.5 off your team's desk after paying for its mistakes. "Right overall" counts every decision, including the ones the tool would hand to a person. "Right when it decided alone" is the number that matters for trust.
| AI tool | Right overall | How often it decided alone | Right when it decided alone | Work saved per 100 decisions |
|---|---|---|---|---|
| J33-Jev, no extra layer | 75.2% | 39.9% | 92.7% | +22.5 |
| Laya (practice-trained), re-tuned | 76.9% | 17.7% | 99.4% | +17.1 |
| Jev (online service) | 73.7% | 37.6% | 89.1% | +13.0 |
| J33-Jev, bigger extra layer | 72.2% | 38.1% | 89.0% | +12.9 |
| J33-Jev, first version, refitted | 71.5% | 34.1% | 87.4% | +8.3 |
| Laya (practice-trained), as delivered | 76.9% | 8.2% | 99.4% | +7.9 |
| J33-Jev, also trained on images | 65.1% | 9.8% | 93.3% | +5.9 |
| J33-Jev, first version | 71.5% | 10.2% | 88.7% | +3.3 |
| Julia-1 | 73.0% | 90.3% | 75.3% | −43.5 |
Not every tool here had the same head start. Every J33-Jev version and both Laya rows studied this test's practice cases. Jev was used exactly as delivered, with no training on them. Julia-1's makers don't publish its training data, and their files mention the same four kinds of work. The test's own notes say trained and untrained scores should not be compared directly. The fair comparison for J33-Jev is Laya.
We also adjusted the confidence setting only for our own model and for Laya (explained below). Jev's and Julia-1's scores were used exactly as delivered.
Confidence that matched accuracy saved more work
Laya was right more often than J33-Jev: 76.9% against 75.2%. When it decided alone, it was right 99.4% of the time and made only 2 mistakes in 2,000 decisions. J33-Jev made 58.
J33-Jev still saved more work, because it decided alone more than twice as often. After we fitted it on held-out practice cases, its confidence was the closest match to its accuracy of any tool here. On average, how sure it said it was came within 1.9 percentage points of how often it was right. Jev came within 3.3 points, with no training or tuning on this test at all. Re-tuned Laya was 13.7 points out.
Laya's most confident 39.7% of decisions were 95% right taken together, yet it decided alone on only 17.7%. The difference, about one decision in five, went to the team.
Part of this is down to how we ran each model. We fitted J33-Jev's confidence setting ourselves, on 120 practice cases it had not seen. For Laya we used the setting stored in its checkpoint. Its documentation says to fit that setting on your own data, and we have not done that for Laya.
Which you prefer depends on your business. If a mistake costs you far more than five right answers save, Laya's careful profile may suit you better. On your real cases, with answers your own staff agree on, the numbers will be different.
Refitting confidence alone more than doubled the work saved
Our first version and its refitted copy gave exactly the same answers. The only difference was the confidence setting. Refitting it raised the share of decisions handled alone from 10.2% to 34.1%, and work saved from +3.3 to +8.3.
A model that undersells itself sends easy work back to your team, as we saw with Laya last time.
Simpler beat bigger, in our one run of each
In our single run of each version, the simplest, with no extra layer, did best. The bigger extra layer was less accurate (72.2%) and saved little more than half as much work (+12.9). Add the bigger layer only if your own tests show it helps.
One version was also trained on photo questions. In our own training-time check it was right 96.1% of the time on those, but on the business cases it scored 65.1% and saved +5.9. It was trained and tuned differently from the others, so we draw no conclusion from it.
Where this pays off
The test covered four kinds of everyday work:
- Customer support: sending each ticket to the right team and spotting the urgent ones.
- Accounts payable: checking whether an invoice is a duplicate or needs a closer look.
- Security: sorting alerts into "needs someone tonight" and "can wait".
- AI assistants: checking whether an automated assistant did its job or needs a person.
J33-Jev was strongest on invoices, right 80.2% of the time. Against Laya, the fair comparison, it trailed on every kind of work: by 0.4 points on invoices and by up to 3.2 points on customer service. A model trained on your own work can still be uneven, so check each kind of work separately.
Take a team handling 10,000 of these decisions a month. This is an example volume, scaled from the rates we measured:
- J33-Jev would handle about 3,985 alone, with around 290 mistakes to fix. Net, about 2,245 decisions off the team's desk.
- Laya, re-tuned would handle about 1,765 alone, with around 10 mistakes. Net, about 1,705.
- Jev would handle about 3,755 alone, with around 410 mistakes. Net, about 1,295, with no training on this test's practice cases.
- Julia-1 would handle about 9,030 alone, but around 2,230 would be wrong. That creates more work than it saves.
Everything not handled alone goes to the team as it does today.
On this test, building our own took about 950 more decisions a month off the desk than Jev. But ours had studied the test's practice cases and Jev had not.
What does it need to run?
Training
We trained our best version on one NVIDIA RTX 4090, a 24 GB graphics card sold for gaming PCs, not a data-centre server. Training took about 10 minutes: four passes over the practice cases at roughly two and a half minutes each. Our other versions were trained on an NVIDIA A40, a 48 GB data-centre card.
Inference (using the trained model)
We ran our best version on the same RTX 4090. It answered each case, five decisions, in about 33 thousandths of a second. That is a different machine from the one in our last test, so the speeds are not directly comparable with the other tools.
It also runs without a GPU. We ran the same model on the processor of an Apple M3 Max laptop, with the graphics chip unused, over all 400 test cases. The answers were identical: 75.2% right and +22.5 work saved per 100 decisions. It took just under a second per case, about 28 times slower than the RTX 4090. For the 10,000 decisions a month in our example, about 2,000 cases, that is roughly half an hour of computer time. A processor is enough for work that can wait a second or two. Anything that must answer instantly, such as a live chat, needs a GPU.
The model file is about 1.09 GB, and you can download it from Hugging Face. The model card shows how to load it in Python and run it on a processor. Google designed EmbeddingGemma 2, the base we built on, to run on phones and laptops.
When is it worth building your own AI decision model?
Here is how we read these results for a business deciding whether to build.
- Test an off-the-shelf tool on your own past cases first. Jev saved +13.0 per 100 decisions with no training on this test's practice cases. If your data may leave your company and that is enough, you may not need to build anything.
- Build when you have the past decisions to train on. Training on your own kind of cases gave the biggest single gain in accuracy we have seen across both articles. Laya's general version was right 36.1% of the time; the version trained on practice cases, 76.9%.
- Work out what a mistake costs you first. It sets how sure the model must be, and it decides whether J33-Jev's profile or Laya's suits you.
- Fit the confidence setting (calibration) on cases kept aside. Refitting it alone took our first version from +3.3 to +8.3, with no change to its answers.
- Start small. In our single run of each version, the simplest beat the bigger one. Add complexity only when your own tests show it helps.
- Keep people in the loop and keep checking. Even the best version here handed six decisions in ten to a person. Plan how those reach your team, and re-test on recent cases every few months.
Frequently asked questions
What is J33-Jev?
J33-Jev is a small System One decision model that J33 built on Google DeepMind's EmbeddingGemma 2. You give it a case and a few fixed questions, and it picks an answer to each and says how sure it is. It has 271.6 million learned values and is free to download from Hugging Face under Apache 2.0.
How did J33-Jev compare with Laya and Jev?
On 2,000 test decisions, J33-Jev saved 22.5 pieces of work per 100 decisions, against 17.1 for a re-tuned Laya from ConvAI Innovations and 13.0 for TypeSafe's Jev. Laya was more accurate (76.9% against 75.2%) and made far fewer mistakes. J33-Jev and Laya had both studied practice cases from the same test; Jev was used as delivered.
Does J33-Jev need a GPU?
Not to run it. On an Apple M3 Max laptop processor, with the graphics chip unused, it gave the same answers in just under a second per case. We trained it in about 10 minutes on one NVIDIA RTX 4090 gaming graphics card, where it answered each case in about 33 thousandths of a second.
When is it worth building your own AI decision model?
When you have a record of past decisions, with the answers your staff settled on, to train it on. Test an off-the-shelf tool on your own cases first, work out what a mistake costs you, and fit the model's confidence on cases kept aside.
How much data do you need to build your own AI decision model?
We trained J33-Jev on 1,080 practice cases, about 5,400 answered questions, and kept 120 cases aside to fit its confidence. We have not tested how few would do. A record of past decisions, with the answers your staff settled on, is the raw material.
How do I download and run J33-Jev?
It is free to download from Hugging Face at huggingface.co/J33-AI/j33-jev under the Apache 2.0 licence. The model card shows how to load it in Python. It runs on an ordinary laptop processor in just under a second per case, or much faster on a GPU.
Working with J33
Our automation and AI agents work starts by working out what a mistake costs your business and testing off-the-shelf tools on your own past cases. If none is good enough, we train a small model on your past decisions and fit its confidence on cases it has never seen. Then we set it up so your people only see the decisions that need them.
How many past decisions does your team have on record?
Tell us which kind takes up the most time and roughly how many you have, with the answers your staff settled on. We will tell you whether an off-the-shelf tool is enough or whether building your own is worth it.
Contact Us