System One AI Models: Which Ones Can You Use in Production?

System One AI Models: Which Ones Can You Use in Production?
System One ModelsAI in ProductionAI AutomationCase StudyOctober 4, 202610 min readBy J33 Tech Team

In short

We tested seven AI tools on 2,000 everyday business decisions. Two of them were right about 73 times in 100. Even so, one saved our imaginary team work and the other made more work for it. The one that saved work knew when it wasn't sure, and handed those cases to a person. The best result was right 99 times out of 100 when it decided alone, and it runs on an ordinary server with no GPU.

What is a System One model? An AI built for one kind of job: you give it a ticket or an invoice and a few fixed questions, and it picks an answer to each and says how sure it is. A general LLM like ChatGPT writes its reply in words instead, which is slower and doesn't come with a built-in measure of how sure it is.

What we tested, at a glance

ModelMade bySizeLicenceNeeds a GPU?Good straight out of the box?How close to real work?
LayaConvAI Innovations421 millionApache 2.0 (free, commercial use allowed)No, ran on an ordinary serverThe general version, no (36% right). The version its makers trained on this test's practice cases, yesBest result (+17.1), but it had studied practice cases from this same test. Promising; test it on your own cases
JevTypeSafeNot publishedOnline service, not downloadableNo, it runs on TypeSafe's serversYesStrong (+13.0). The most honest about its own doubt. Your data has to leave your company
Julia-1Supersonic Labs144 millionApache 2.0No, ran on an ordinary serverAccurate (73%), but its confidence scores were not a good guide to when it was rightNot for deciding alone with its scores as delivered (−43.5)
GLiNER2.5-DecideFastino486 millionApache 2.0No, ran on an ordinary serverClose to guessing on these cases (51% right), but it knew it and passed almost everything to a personNot yet (+0.5)
CLMContrastive-LM8 billionApache 2.0We only ran it with a GPUNot meant to be; its makers say to train it first. 41% right before, 63% after we trained itNot yet (+2.6 after training)
levInterfaze4 billionApache 2.0Yes, it would not run without an NVIDIA GPUHonest about its doubt, but not accurate enough to clear our bar oftenBreak-even (−1.6)
Qwen3, for comparison (a general chat AI, not built for this job)Alibaba (Qwen team)8 billionApache 2.0We only ran it with a GPUReasonable answers (62% right), but when simply asked how sure it was, it always said at least 70%Not on its own for deciding when to act (−120.2)

Size is how many learned values the model has; bigger models usually need more powerful hardware. The number in brackets is the work each tool saved per 100 decisions, explained below.

Think of AI as a new hire

Picture a new person joining your support team. In their first week, you don't let them handle every customer alone. You let them deal with the easy cases and ask a colleague about the rest.

What makes a good new hire isn't just being right most of the time. It's knowing when they might be wrong, and asking.

AI tools that make business decisions work the same way. Which team should get this support ticket? Is this invoice a duplicate? Does this security alert need someone tonight? Since mid-September 2026, a wave of AI tools has appeared that answer questions like these. These are the System One models, a format started by TypeSafe's Jev.

Each one comes with its own test scores, and those are hard to compare. We wanted to know something more useful: how much work can you safely hand to each one?

What we did

We used a public set of practice cases, typed-decisions. It has 400 test cases from four kinds of work: customer service, invoices, security alerts and checking on AI assistants. Each case, such as one support ticket, comes with five questions, like which team should handle it and how urgent it is. That makes 2,000 decisions. The cases are made up for testing, and the "right answers" were set by another AI model rather than by people. So "right" here means "agreed with that model". It agrees with itself only about 73 times in 100, so a score well above that partly means a tool has learned the labeller's habits.

We agreed on a score before we looked at any results. This was the most important step.

  • If the AI decides alone and gets it right, it saves a person some work. We counted that as +1.
  • If it decides alone and gets it wrong, someone has to spot the mistake, fix it and deal with any harm. We counted that as −5.
  • If it hands the case to a person, nothing changes from today. That counts as 0.

So one mistake wipes out the good done by five right answers. That means the AI should only decide alone when it is more than 83% sure. Below that, its mistakes cost more than its right answers save, and it's better to ask a person.

Your numbers may be different. If a mistake costs you ten times what a right answer saves, the AI needs to be about 91% sure before it decides alone.

We tested seven tools, listed in the table above. Five are open models anyone can download. Jev is an online service. Qwen3 is a well-known general chat AI, which we added to see what happens if you simply ask a chat AI to make these decisions. We ran all the downloadable tools on the same computer. Jev runs on TypeSafe's servers, so we sent it the same cases over the internet.

We also tried teaching one tool. We showed CLM past cases along with the right answers, to see how fast it would learn.

What we found

The last column is the one that matters. One unit is one decision a person no longer has to make. So +17.1 means that for every 100 decisions, the tool took about 17 off your team's desk after paying for its mistakes. A minus number means it created more work than it saved.

"Right overall" counts every decision, including the ones the tool would hand to a person. "Right when it decided alone" is the number that matters for trust.

AI toolRight overallHow often it decided aloneRight when it decided aloneWork saved per 100 decisions
Laya (trained on practice cases), setting re-tuned76.9%17.7%99.4%+17.1
Jev (online service)73.7%37.6%89.1%+13.0
Laya (trained on practice cases), as delivered76.9%8.2%99.4%+7.9
CLM, after teaching it63.0%5.9%90.7%+2.6
GLiNER2.5-Decide50.6%0.8%93.3%+0.5
lev63.4%21.8%82.1%−1.6
Laya, general version36.1%3.6%45.8%−8.1
Julia-173.0%90.3%75.3%−43.5
Qwen3, chat AI (for comparison)62.0%95.3%62.3%−120.2
CLM, before teaching it41.3%63.3%42.8%−154.0

For comparison, just picking the most common answer every time was right 47.9% of the time.

Work saved per 100 decisions, after the cost of mistakes
Work saved per 100 decisions, after the cost of mistakes. Blue saves work, orange creates extra work.

Being right often isn't enough

Jev and Julia-1 were right about equally often: 73.7% and 73.0% of the time. You might expect them to be equally useful. They weren't. Jev saved 13 pieces of work per 100 decisions. Julia-1 created 43.5 pieces of extra work.

The reason is how well each one's confidence matched reality. When Jev said it was about 96% sure, it was right 92% of the time. Its confidence meant something.

Julia-1 gave almost every answer a confidence above 83%, so it decided alone nearly every time, and it was right on about three in four of those. We used its scores as delivered, as we did for every tool except Laya. Its confidence doesn't separate its right answers from its wrong ones well enough to hand it much work. Julia-1's makers don't publish its training data, and their files mention these same four kinds of work, so it may also have had a head start on this test.

If we had chosen by the headline score alone, we would have picked a tool that creates more work than it saves. That is the main lesson from this test.

The best result came with two catches

The best result came from Laya. But this was not the general version you would download for any job. Its makers had trained it on the practice half of this same public case set, so it had a head start the other tools didn't. It scored 76.9%, above the 73% at which the labeller agrees with itself, so part of that is learning the labeller's habits. Laya's general version scored 36% and created more work than it saved.

The second catch was a confidence setting. Laya's own documentation says to re-tune it on your own data before trusting its scores. Left as delivered, it decided alone on only 8% of decisions. When we re-tuned it, it decided alone on 18%, was still right 99.4% of the time, and saved more than twice as much work.

The instruction is in the documentation, but it is easy to skip. You only see what skipping it costs by testing the tool on your own cases.

A chat AI gave decent answers, but couldn't rate its own doubt when simply asked

Qwen3 did well for a general chat AI that was never built for this job. It was right 62% of the time, more often than some of the specialist tools.

The weak spot was confidence. We asked it the simple way, to write how sure it was next to each answer, and it said at least 70% every time. A tool that always sounds sure can't tell you when to ask a person. We only tried that simple way; other ways of getting a confidence out of a chat AI might do better. In our setup it was also about 250 times slower than the fastest specialist tool.

A little teaching goes a long way

CLM's makers say it is meant to be trained before use, and the results agree. Before teaching, it did worse than just picking the most common answer. After we showed it 50 past cases with the right answers (plus 120 more we used to check its progress), it was right 56% of the time. After 300, it was right 62% of the time. Many teams already have that many solved cases sitting in their old tickets.

Where this pays off

The test covered four kinds of everyday work, and each is a real cost for a business today:

  • Customer support: sending each ticket to the right team, spotting the urgent ones and deciding which need a person.
  • Accounts payable: checking whether an invoice is a duplicate or needs a closer look before it is paid.
  • Security: sorting alerts into "needs someone tonight" and "can wait until morning".
  • AI assistants: checking whether an automated assistant did its job properly or needs a person to step in.

Here is what the difference between tools looks like on a desk. Take a team handling 10,000 of these decisions a month (an example volume, using the rates we measured):

  • With Laya, about 1,770 decisions a month would be handled without anyone touching them, with around 10 mistakes to catch and fix. Everything else goes to the team as it does today.
  • With Julia-1, about 9,000 decisions would be handled alone, but around 2,200 of them would be wrong. That is more clean-up work than it saves.

Both are right roughly three times in four overall. One gives your team time back; the other gives them a mess to clear up. Finding out which is which, on your own cases, is the first thing we do in a client project.

Does the best tool need a GPU?

No. AI tools often run on a GPU. We ran the three smallest tools again without one: on an ordinary server with 8 processor cores, and on a high-end laptop (an Apple M3 Max).

AI toolWith a GPUOrdinary serverHigh-end laptopDecisions per hour on the server
Julia-10.01 seconds0.6 seconds0.4 secondsabout 28,000
Laya (the best result)0.05 seconds2.6 seconds1.3 secondsabout 6,800
GLiNER0.07 seconds1.5 seconds1.0 secondsabout 12,000

Times are for one case, which is five decisions.

The answers were almost identical. On the ordinary server, Laya was right 76.6% of the time, against 76.9% with a GPU. When it decided alone, it was still right 99.4% of the time, and it saved the same +17.1 per 100 decisions.

At about 6,800 decisions an hour, one ordinary server can handle around 160,000 decisions a day. Many support, accounts and security teams handle far fewer than that. A GPU made these tools roughly 20 to 55 times faster. That helps if you need instant answers at huge volumes, but it didn't make the answers any better.

Two warnings. A very small server is too slow: on one with only 2 processor cores, Laya took about 12 seconds per case, and up to 20 on the longest ones. And we only ran the bigger tools (CLM, lev and the chat AI) with a GPU; lev would not run at all without an NVIDIA one.

What we recommend

If you're thinking about letting AI make decisions in your business, here's the order we'd go in.

  1. Work out what a mistake costs you. This tells you how sure the AI must be before it decides alone. It can also change which tool is best for you.
  2. Don't choose by the headline score. Ask how much of your work the tool can safely take on, and how often it's right when it does.
  3. Test it on your own past cases. A few hundred cases is enough to see whether you can trust how sure the tool says it is. It will also catch problems like Laya's confidence setting.
  4. Expect to teach it a little. The best result here came from a tool trained on practice cases from this same test. A few hundred examples made a big difference to CLM.
  5. Choose where it runs based on your data rules. If your data must stay inside your company, the best tool here runs on an ordinary server you may already have. If your data can go outside, Jev worked well without any changes and was the most honest about how sure it was.
  6. Keep people in the loop. Even with the best tool, only its most confident 40% of decisions were 95% right; beyond that, mistakes climb. Plan how the rest reach your team, and how quickly.
  7. Keep checking. Your business changes, and a tool that worked well last quarter can slowly get worse. Test it on recent cases every few months.

Working with J33

This is the work we do for clients. We start by working out what a mistake costs your business. Then we test the possible tools on your own past cases, train the best one, and set it up so your people only see the decisions that need them.

Does your team sort tickets, check invoices or deal with alerts every day?

Tell us which decisions take up the most time. We will tell you how we would test whether AI can safely take them on.

Contact Us