In short
We tested seven AI tools on 2,000 everyday business decisions. Two of them were right about 73 times in 100. Even so, one saved our imaginary team work and the other made more work for it. The one that saved work knew when it wasn't sure, and handed those cases to a person. The best result was right 99 times out of 100 when it decided alone, and it runs on an ordinary server with no GPU.
What we tested, at a glance
| Model | Made by | Size | Licence | Needs a GPU? | Good straight out of the box? | How close to real work? |
|---|---|---|---|---|---|---|
| Laya | ConvAI Innovations | 421 million | Apache 2.0 (free, commercial use allowed) | No, ran on an ordinary server | The general version, no (36% right). The version its makers trained on this test's practice cases, yes | Best result (+17.1), but it had studied practice cases from this same test. Promising; test it on your own cases |
| Jev | TypeSafe | Not published | Online service, not downloadable | No, it runs on TypeSafe's servers | Yes | Strong (+13.0). The most honest about its own doubt. Your data has to leave your company |
| Julia-1 | Supersonic Labs | 144 million | Apache 2.0 | No, ran on an ordinary server | Accurate (73%), but its confidence scores were not a good guide to when it was right | Not for deciding alone with its scores as delivered (−43.5) |
| GLiNER2.5-Decide | Fastino | 486 million | Apache 2.0 | No, ran on an ordinary server | Close to guessing on these cases (51% right), but it knew it and passed almost everything to a person | Not yet (+0.5) |
| CLM | Contrastive-LM | 8 billion | Apache 2.0 | We only ran it with a GPU | Not meant to be; its makers say to train it first. 41% right before, 63% after we trained it | Not yet (+2.6 after training) |
| lev | Interfaze | 4 billion | Apache 2.0 | Yes, it would not run without an NVIDIA GPU | Honest about its doubt, but not accurate enough to clear our bar often | Break-even (−1.6) |
| Qwen3, for comparison (a general chat AI, not built for this job) | Alibaba (Qwen team) | 8 billion | Apache 2.0 | We only ran it with a GPU | Reasonable answers (62% right), but when simply asked how sure it was, it always said at least 70% | Not on its own for deciding when to act (−120.2) |
Size is how many learned values the model has; bigger models usually need more powerful hardware. The number in brackets is the work each tool saved per 100 decisions, explained below.
Think of AI as a new hire
Picture a new person joining your support team. In their first week, you don't let them handle every customer alone. You let them deal with the easy cases and ask a colleague about the rest.
What makes a good new hire isn't just being right most of the time. It's knowing when they might be wrong, and asking.
AI tools that make business decisions work the same way. Which team should get this support ticket? Is this invoice a duplicate? Does this security alert need someone tonight? Since mid-September 2026, a wave of AI tools has appeared that answer questions like these. These are the System One models, a format started by TypeSafe's Jev.
Each one comes with its own test scores, and those are hard to compare. We wanted to know something more useful: how much work can you safely hand to each one?
What we did
We used a public set of practice cases, typed-decisions. It has 400 test cases from four kinds of work: customer service, invoices, security alerts and checking on AI assistants. Each case, such as one support ticket, comes with five questions, like which team should handle it and how urgent it is. That makes 2,000 decisions. The cases are made up for testing, and the "right answers" were set by another AI model rather than by people. So "right" here means "agreed with that model". It agrees with itself only about 73 times in 100, so a score well above that partly means a tool has learned the labeller's habits.
We agreed on a score before we looked at any results. This was the most important step.
- If the AI decides alone and gets it right, it saves a person some work. We counted that as +1.
- If it decides alone and gets it wrong, someone has to spot the mistake, fix it and deal with any harm. We counted that as −5.
- If it hands the case to a person, nothing changes from today. That counts as 0.
So one mistake wipes out the good done by five right answers. That means the AI should only decide alone when it is more than 83% sure. Below that, its mistakes cost more than its right answers save, and it's better to ask a person.
Your numbers may be different. If a mistake costs you ten times what a right answer saves, the AI needs to be about 91% sure before it decides alone.
We tested seven tools, listed in the table above. Five are open models anyone can download. Jev is an online service. Qwen3 is a well-known general chat AI, which we added to see what happens if you simply ask a chat AI to make these decisions. We ran all the downloadable tools on the same computer. Jev runs on TypeSafe's servers, so we sent it the same cases over the internet.
We also tried teaching one tool. We showed CLM past cases along with the right answers, to see how fast it would learn.
What we found
The last column is the one that matters. One unit is one decision a person no longer has to make. So +17.1 means that for every 100 decisions, the tool took about 17 off your team's desk after paying for its mistakes. A minus number means it created more work than it saved.
"Right overall" counts every decision, including the ones the tool would hand to a person. "Right when it decided alone" is the number that matters for trust.
| AI tool | Right overall | How often it decided alone | Right when it decided alone | Work saved per 100 decisions |
|---|---|---|---|---|
| Laya (trained on practice cases), setting re-tuned | 76.9% | 17.7% | 99.4% | +17.1 |
| Jev (online service) | 73.7% | 37.6% | 89.1% | +13.0 |
| Laya (trained on practice cases), as delivered | 76.9% | 8.2% | 99.4% | +7.9 |
| CLM, after teaching it | 63.0% | 5.9% | 90.7% | +2.6 |
| GLiNER2.5-Decide | 50.6% | 0.8% | 93.3% | +0.5 |
| lev | 63.4% | 21.8% | 82.1% | −1.6 |
| Laya, general version | 36.1% | 3.6% | 45.8% | −8.1 |
| Julia-1 | 73.0% | 90.3% | 75.3% | −43.5 |
| Qwen3, chat AI (for comparison) | 62.0% | 95.3% | 62.3% | −120.2 |
| CLM, before teaching it | 41.3% | 63.3% | 42.8% | −154.0 |
For comparison, just picking the most common answer every time was right 47.9% of the time.
Being right often isn't enough
Jev and Julia-1 were right about equally often: 73.7% and 73.0% of the time. You might expect them to be equally useful. They weren't. Jev saved 13 pieces of work per 100 decisions. Julia-1 created 43.5 pieces of extra work.
The reason is how well each one's confidence matched reality. When Jev said it was about 96% sure, it was right 92% of the time. Its confidence meant something.
Julia-1 gave almost every answer a confidence above 83%, so it decided alone nearly every time, and it was right on about three in four of those. We used its scores as delivered, as we did for every tool except Laya. Its confidence doesn't separate its right answers from its wrong ones well enough to hand it much work. Julia-1's makers don't publish its training data, and their files mention these same four kinds of work, so it may also have had a head start on this test.
If we had chosen by the headline score alone, we would have picked a tool that creates more work than it saves. That is the main lesson from this test.
The best result came with two catches
The best result came from Laya. But this was not the general version you would download for any job. Its makers had trained it on the practice half of this same public case set, so it had a head start the other tools didn't. It scored 76.9%, above the 73% at which the labeller agrees with itself, so part of that is learning the labeller's habits. Laya's general version scored 36% and created more work than it saved.
The second catch was a confidence setting. Laya's own documentation says to re-tune it on your own data before trusting its scores. Left as delivered, it decided alone on only 8% of decisions. When we re-tuned it, it decided alone on 18%, was still right 99.4% of the time, and saved more than twice as much work.
The instruction is in the documentation, but it is easy to skip. You only see what skipping it costs by testing the tool on your own cases.
A chat AI gave decent answers, but couldn't rate its own doubt when simply asked
Qwen3 did well for a general chat AI that was never built for this job. It was right 62% of the time, more often than some of the specialist tools.
The weak spot was confidence. We asked it the simple way, to write how sure it was next to each answer, and it said at least 70% every time. A tool that always sounds sure can't tell you when to ask a person. We only tried that simple way; other ways of getting a confidence out of a chat AI might do better. In our setup it was also about 250 times slower than the fastest specialist tool.
A little teaching goes a long way
CLM's makers say it is meant to be trained before use, and the results agree. Before teaching, it did worse than just picking the most common answer. After we showed it 50 past cases with the right answers (plus 120 more we used to check its progress), it was right 56% of the time. After 300, it was right 62% of the time. Many teams already have that many solved cases sitting in their old tickets.
Where this pays off
The test covered four kinds of everyday work, and each is a real cost for a business today:
- Customer support: sending each ticket to the right team, spotting the urgent ones and deciding which need a person.
- Accounts payable: checking whether an invoice is a duplicate or needs a closer look before it is paid.
- Security: sorting alerts into "needs someone tonight" and "can wait until morning".
- AI assistants: checking whether an automated assistant did its job properly or needs a person to step in.
Here is what the difference between tools looks like on a desk. Take a team handling 10,000 of these decisions a month (an example volume, using the rates we measured):
- With Laya, about 1,770 decisions a month would be handled without anyone touching them, with around 10 mistakes to catch and fix. Everything else goes to the team as it does today.
- With Julia-1, about 9,000 decisions would be handled alone, but around 2,200 of them would be wrong. That is more clean-up work than it saves.
Both are right roughly three times in four overall. One gives your team time back; the other gives them a mess to clear up. Finding out which is which, on your own cases, is the first thing we do in a client project.
Does the best tool need a GPU?
No. AI tools often run on a GPU. We ran the three smallest tools again without one: on an ordinary server with 8 processor cores, and on a high-end laptop (an Apple M3 Max).
| AI tool | With a GPU | Ordinary server | High-end laptop | Decisions per hour on the server |
|---|---|---|---|---|
| Julia-1 | 0.01 seconds | 0.6 seconds | 0.4 seconds | about 28,000 |
| Laya (the best result) | 0.05 seconds | 2.6 seconds | 1.3 seconds | about 6,800 |
| GLiNER | 0.07 seconds | 1.5 seconds | 1.0 seconds | about 12,000 |
Times are for one case, which is five decisions.
The answers were almost identical. On the ordinary server, Laya was right 76.6% of the time, against 76.9% with a GPU. When it decided alone, it was still right 99.4% of the time, and it saved the same +17.1 per 100 decisions.
At about 6,800 decisions an hour, one ordinary server can handle around 160,000 decisions a day. Many support, accounts and security teams handle far fewer than that. A GPU made these tools roughly 20 to 55 times faster. That helps if you need instant answers at huge volumes, but it didn't make the answers any better.
Two warnings. A very small server is too slow: on one with only 2 processor cores, Laya took about 12 seconds per case, and up to 20 on the longest ones. And we only ran the bigger tools (CLM, lev and the chat AI) with a GPU; lev would not run at all without an NVIDIA one.
What we recommend
If you're thinking about letting AI make decisions in your business, here's the order we'd go in.
- Work out what a mistake costs you. This tells you how sure the AI must be before it decides alone. It can also change which tool is best for you.
- Don't choose by the headline score. Ask how much of your work the tool can safely take on, and how often it's right when it does.
- Test it on your own past cases. A few hundred cases is enough to see whether you can trust how sure the tool says it is. It will also catch problems like Laya's confidence setting.
- Expect to teach it a little. The best result here came from a tool trained on practice cases from this same test. A few hundred examples made a big difference to CLM.
- Choose where it runs based on your data rules. If your data must stay inside your company, the best tool here runs on an ordinary server you may already have. If your data can go outside, Jev worked well without any changes and was the most honest about how sure it was.
- Keep people in the loop. Even with the best tool, only its most confident 40% of decisions were 95% right; beyond that, mistakes climb. Plan how the rest reach your team, and how quickly.
- Keep checking. Your business changes, and a tool that worked well last quarter can slowly get worse. Test it on recent cases every few months.
Working with J33
This is the work we do for clients. We start by working out what a mistake costs your business. Then we test the possible tools on your own past cases, train the best one, and set it up so your people only see the decisions that need them.
Does your team sort tickets, check invoices or deal with alerts every day?
Tell us which decisions take up the most time. We will tell you how we would test whether AI can safely take them on.
Contact Us