> ## Content Index
> Fetch the complete content index at: https://blog.uniportal.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Can Fast AI Models like Jev Triage MSP Work? Four Models Tested on Tickets, Alerts and Emails
- URL: https://blog.uniportal.ai/can-fast-ai-models-like-jev-triage-msp-work-four-models-tested-on-tickets-alerts-and-emails/
- Published: 2026-09-23T20:44:45.000Z
- Updated: 2026-09-23T20:44:44.000Z
- Author: Ricky Sohal

A lot of service-desk work starts with small decisions. What is this ticket about? How urgent is it? Does this RMM alert need someone now? Is this email a new request, a reply, or something to ignore?

These are classification tasks, and small, fast AI models are often suggested for exactly this kind of work. General-purpose leaderboards don't show how well they'd do it, though. They measure things like coding and maths, not whether a model can tell an outage from an urgently worded note about a printer.

This study is part of Uniportal's ongoing research into how AI models perform on MSP operations. We compared four fast models across three everyday MSP classification tasks to understand where they perform well, where they differ, and what that means for anyone considering AI in service delivery. 

## At a glance

- **All four models are accurate on well-defined classification.** Ticket category, a security-incident flag and RMM alert triage all scored 95–99%. No model labelled an urgent alert as noise.
- **Priority is the hardest call, for models and people alike.** Two expert reviewers agreed with each other on 80% of tickets. Three of the four models reached 87–88% against the final answer key.
- **Each model has a consistent lean** when it gets priority wrong: some towards more urgent, some towards less.
- **Confidence scores differ a lot in how trustworthy they are**, and that decides how safely a model can work without a human check.
- **Speed and cost vary far more than accuracy.** Median response times ranged from 0.14s to 1.7s. Cost per 10,000 tickets ranged from $0.50 to $14.64.

## What we tested

**Three tasks drawn from day-to-day MSP operations:**

1. **PSA ticket triage.** Assign a category (10 options, from account access to server infrastructure), a priority (P1–P4 by business impact and urgency), and a flag for possible compromise, even when the requester doesn't mention it.
2. **RMM alert triage.** Decide whether an alert needs a technician now, can wait for the next business day, or needs no action. The last group includes things like a workstation offline after hours, a CPU spike inside the patch window, or an alert that has already cleared.
3. **Helpdesk inbox sorting.** Label each email as a new request, an update to an existing ticket, an auto-reply, a vendor notice, sales or marketing, or phishing.

**Four fast models, each at its quickest configuration:**

- **Jev 1.13** (TypeSafe). A decision model that returns typed answers and probabilities rather than text.
- **Gemini 3.8 Flash** (Google), with thinking set as low as the API allows.
- **GPT-5.4 mini** (OpenAI), with reasoning turned off.
- **Claude Haiku 4.5** (Anthropic).

## How we built the test

The answer key matters more than anything else in a benchmark, so most of the effort went there.

**Answers first, then text.** For each of 1,080 items, we fixed the correct answer before any text was written, then had an AI model write a realistic ticket, alert or email to match. About a third were deliberately tricky. Examples include an "URGENT!!!" email about a cosmetic issue, a calm note describing a site-wide outage, a phishing email styled as a Microsoft notice, and a disk alert where only the device's role shows whether it matters.

**Three writers, none of them under test.** GPT-5.6, Gemini Pro and Claude Fable each wrote a third of the items, so no single writing style dominates.

**Two independent reviewers on every item.** Each item was also labelled by the two writer models that didn't produce it, without seeing the intended answer. The final answer is the majority of the three. The 19 items with no majority were removed, leaving 1,061.

**Equal tuning, then a held-out test.** All four models started with identical instructions. On the ticket task, each tried the same two refined prompts on a separate tuning set and kept whichever performed best. The 795 items used for the final results were never used during tuning.

**Repeated runs and real costs.** Every test ran three times to measure consistency. Costs come from the tokens each model actually used, at published list prices in September 2026.

## Result 1: Well-defined classification is a solved problem for fast models

![](https://blog.uniportal.ai/content/images/2026/09/01-accuracy-by-task.png)

Accuracy on 795 held-out MSP items

Each dot is a model's accuracy. Each line shows the range the true score likely falls in (a 95% confidence interval). Where lines overlap, the difference between models isn't meaningful.

Ticket category, the security flag and RMM alert triage all fall between 95% and 99%. Gemini 3.8 Flash was the most accurate on ticket category, at 98.6%.

Across 225 RMM alerts, no model labelled an urgent alert as noise. Gemini 3.8 Flash and Claude Haiku 4.5 identified every urgent alert. Jev and GPT-5.4 mini each rated one as "next business day".

**What this means:** for routing and categorising, the four fast models tested all perform well. The model choice here comes down to other factors, such as speed, cost and consistency.

## Result 2: Priority, and why it's the hardest call

Priority depends on judgement. Is one user's laptop freezing during dictation a P2 or a P3? It depends on who they are, what they're working on and whether there's a workaround. Our two expert reviewers agreed with each other on only 80% of tickets, which shows how much room for interpretation priority leaves.

Against the final answer key, Gemini 3.8 Flash (88.4%), Jev (86.7%) and GPT-5.4 mini (86.7%) performed within each other's margin of error, and Claude Haiku 4.5 reached 80.1%. When models were wrong, they were almost always off by only one level: 97–98% of answers were within one priority level of the key.

The more useful finding is the direction of each model's mistakes:

![](https://blog.uniportal.ai/content/images/2026/09/07-confidently-wrong.png)

- **Claude Haiku 4.5 leans towards higher priority.** It rated 15.9% of tickets as more urgent than the key.
- **GPT-5.4 mini leans towards lower priority.** It rated 8.7% as less urgent than the key.
- **Jev and Gemini 3.8 Flash lean slightly higher.**

In a service desk, a ticket rated too high costs some attention. A ticket rated too low can leave a client waiting, so it helps to know which way a model leans.

We also looked at 22 routine-looking tickets that each contained a subtle sign of compromise, such as an MFA prompt the user didn't trigger or a supplier asking to change bank details. The models raised the security flag on 20–22 of them, but assigned the key's priority on 12–16\. **A good design practice follows from this: when a ticket is flagged as a possible security incident, let the flag raise its priority as well.**

## Result 3: Two inbox edge cases

![](https://blog.uniportal.ai/content/images/2026/09/06-inbox-traps.png)

**A new issue sent as a reply in an old ticket thread.** Clients often reply to a resolved ticket when they have an unrelated new problem. Gemini 3.8 Flash identified all 16 of these as new requests, and Jev identified 14\. GPT-5.4 mini and Claude Haiku 4.5 each identified 5, treating most of the rest as updates to the old ticket.

**A client forwarding a suspicious email to ask whether it's genuine.** GPT-5.4 mini and Claude Haiku 4.5 labelled all 10 as phishing, and Jev labelled 9\. Gemini 3.8 Flash labelled all 10 as new requests, reading them as a client asking for help. Both readings are reasonable, and our own reviewers were divided on these items as well. On phishing sent directly to the helpdesk, the models caught 24–26 of 27.

**What this means:** some classifications depend on an organisation's own process, not just the content of the email. Stating that process in the model's instructions, for example "forwarded suspicious emails go to the security queue", removes the ambiguity.

## Result 4: Confidence you can act on

Every model in this study reports how confident it is. That confidence matters in practice: a common pattern is to act automatically on high-confidence answers and send low-confidence ones to a technician. The pattern works best when a model's mistakes come with lower confidence.

![](https://blog.uniportal.ai/content/images/2026/09/02-priority-error-direction.png)

When Jev got priority wrong, it reported 90% or higher confidence on 9 of 46 mistakes. Claude Haiku 4.5 did so on 26 of 69, Gemini 3.8 Flash on 22 of 40, and GPT-5.4 mini on 36 of 46\. Jev returns a calculated probability. The other three models state their confidence as part of their answer, which helps explain the difference.

**What this means:** when deciding how much to automate, how reliable the confidence score is matters as much as accuracy.

## Result 5: Response time

![](https://blog.uniportal.ai/content/images/2026/09/04-latency.png)

Time to classify one ticket

We sent 100 tickets one at a time and timed each response. Jev had a median of 144ms, GPT-5.4 mini 845ms and Claude Haiku 4.5 1.2s. For all three, the slowest 5% of requests finished in around 2 seconds or less. Gemini 3.8 Flash had a median of 1.7s, and its slowest 5% took about 9 seconds or longer.

**What this means:** for background work, such as tagging tickets as they arrive, any of these response times is fine. For anything a person is waiting on, it's worth looking at the slowest responses as well as the typical one. These measurements come from a single location on a single day.

## Result 6: Cost

![](https://blog.uniportal.ai/content/images/2026/09/03-cost-vs-accuracy.png)

Accuracy vs Cost

Jev and Gemini 3.8 Flash had the highest average accuracy across all five labels (94.6%). Classifying 10,000 tickets costs $0.50 with Jev, about $8 with Gemini 3.8 Flash or GPT-5.4 mini, and $14.64 with Claude Haiku 4.5.

![](https://blog.uniportal.ai/content/images/2026/09/05-monthly-cost.png)

What classification costs a mid-sized MSP each month

At the volume of a mid-sized MSP (4,000 tickets, 30,000 RMM alerts and 8,000 inbox emails a month), the monthly cost ranges from $1.24 to $31.30\. At that scale, all four models are inexpensive. The differences become significant for platforms running classification across many organisations. Gemini 3.8 Flash's current price is a promotional rate until December 31, 2026.

## Model profiles

**Jev 1.13 (TypeSafe)**

- *Strengths:* Fastest and lowest cost in the study. The most reliable confidence scores. Highest inbox accuracy (94.2%), and it handled old-thread replies well (14 of 16).
- *Considerations:* It caught 24 of 27 direct phishing emails, against 26 for the other three. It returns decisions and probabilities, not written text, and you configure it with typed questions rather than a prompt.

**Gemini 3.8 Flash (Google)**

- *Strengths:* Highest accuracy on ticket category (98.6%) and priority (88.4%). It identified every urgent alert and every old-thread reply.
- *Considerations:* Its response times vary more than the others'. It reads forwarded suspicious emails as client requests unless instructed otherwise.

**GPT-5.4 mini (OpenAI)**

- *Strengths:* Consistent response times (845ms median). Strong phishing detection (26 of 27 direct, 10 of 10 forwarded). The fewest tickets rated more urgent than the key.
- *Considerations:* It leans towards lower priority, and it often reports high confidence on incorrect answers.

**Claude Haiku 4.5 (Anthropic)**

- *Strengths:* The most consistent: identical answers across repeated runs on alerts, inbox and category. Highest RMM alert accuracy (97.8%) and strong phishing detection.
- *Considerations:* It leans towards higher priority, and it had the highest cost per ticket in this study. Its results were noticeably better with native structured output than with a tool call.

## Key takeaways

For anyone applying fast models to MSP classification, the results point to a few practical lessons:

1. **Test on the actual job.** General benchmarks don't predict performance on tickets, alerts and inbox triage. Task-specific tests do.
2. **Measure more than accuracy.** Confidence reliability, response-time consistency, run-to-run consistency and cost all shape how a model behaves in production.
3. **Put the process in the instructions.** Where the right answer depends on how an MSP works, such as how to handle forwarded phishing, stating it explicitly makes model behaviour predictable.
4. **Connect related signals.** A security flag should influence priority, not sit beside it.
5. **Keep people in the loop where confidence is low.** This works best with a model whose confidence scores track its accuracy.

---

*Uniportal gives MSPs an AI teammate that works across Microsoft 365, Google Workspace, your PSA, your RMM, and your documentation with your technician approving every change.* [*Learn more*](https://uniportal.ai/?ref=blog.uniportal.ai)*.*