Burnt and Toast
Studios

From the studio / Jev, part 1 of 3

What use is an AI
that can’t write?

I tested Jev on customer email routing and spam classification, looking at speed, repeatability and whether its confidence matched the results.

A stylised portrait of Leigh beside the words Jev can’t write and Part 1 of 3.

Jev cannot write an email, an article or a politely worded explanation of why a meeting could have been an email. That is quite a list of things to give up in an AI model.

It did, however, make me curious. A lot of useful work involves choosing between a few answers. Which team should handle this message? Does it look like spam? How certain is that decision?

My first Jev video is now live. It follows a weekend spent testing a narrower question: could a model that does not generate prose be useful for this kind of work?

Watch Jev, part one.

Giving it a job it can actually do

TypeSafe AI’s Jev takes information and answers structured questions. You give it the available choices, or ask for a score or a yes/no decision. It returns an answer with a probability. It does not write an explanation of its reasoning.

For the first test, I used customer email routing. With Claude’s help, I wrote 100 messages for a fictional software company and labelled the team that should receive each one.

Nineteen were deliberately awkward. For those, I recorded a second acceptable answer before seeing the results. A request to close an account that also asks about the next bill can reasonably involve more than one team. I wanted the marking to acknowledge that.

Each message went to Jev and GPT-5.4 mini, through the same gateway, one request at a time, alternating which model went first. I repeated the exercise three times.

That gave me something I could inspect: the decision, how long it took and whether the same message received the same answer again.

Speed was only part of the result

Jev was faster in these runs. The size of that advantage belongs to this setup, these requests and this comparison. It is not a promise about the speed of somebody else’s application.

The repeatability interested me too. Across the three passes, Jev assigned every email to the same category each time. A consistent answer can still be wrong, of course. But when a customer message is being routed, changing the destination between runs is something I would want to understand.

The video shows the timing and accuracy results together. A quick decision is only useful if it is good enough for the job.

A confidence score needs checking

The other question was whether Jev’s probabilities were useful. A number next to an answer can look reassuring. It still needs checking against what actually happened.

For this, I used 500 messages from the public SMS Spam Collection, labelled as spam or legitimate messages. These were real messages rather than emails I had written with an AI’s help.

In that sample, Jev reported at least 95% confidence on 428 messages and classified all 428 correctly. The film also shows the message it wrongly flagged. That is evidence from one sample, not proof that its confidence will remain reliable on unfamiliar material.

The collection is also old and well known. It may have appeared in the models’ training data. The small gap between their results is not a sound basis for declaring a winner.

What I would take into the next test

This was one person, one weekend and a few hundred examples. The email labels were mine. The text-message collection has its limitations. I would want to test representative material from a real application before deciding what to automate or where to ask a person to review a decision.

What interests me is having a clearer choice about the work I give a model. If I need a category, I can examine how well it chooses one. If I need an explanation, that is a different requirement, and Jev will not provide it.

Part one is the first check. In part two, I go looking for the weaknesses, including a darts question that causes rather more trouble than it should.

Watch the five-minute investigation.

About the video: the narration uses an AI clone of my own voice, and the presenter is an animated avatar of me. The test results were recorded on 21 September 2026. The video was published on 22 September.

Dataset credit: Almeida, T. and Hidalgo, J. (2011), SMS Spam Collection, UCI Machine Learning Repository.

Read this next

Share this
X LinkedIn Facebook WhatsApp Email

Share to Instagram

Share or save the cover image, then add it to your Instagram story or post. For a story, add a Link sticker with the article address below.

Article cover
Save image

Copied