Most AI Work Isn't Writing, It's Choosing: Testing Jev

In mid-September a San Francisco company, TypeSafe AI, released a new AI model called Jev to a small group of early users. Within days, some of the biggest developer platforms had added it. It still got less public attention than the big chatbot launches. For small and medium businesses, we think it may matter more.
We've been trialling it on our own workflows. Here's what we found, and why we think it opens up something real for Australian SMEs.
Most AI work isn't writing. It's choosing.
When most people think of AI, they think of something that writes, such as ChatGPT and its competitors. Ask a question and it composes an answer word by word, like a conversation.
But when AI is put to work inside a business process, it's rarely asked to hold a conversation. Most of the time it's asked to choose:
- Which account code does this cost belong to?
- Is this document an invoice, a receipt or a statement?
- Does this bill match a purchase order: yes or no?
- Does this need escalating to a person, or can it go straight through?
- Which of these five suppliers is this email from?
- How urgent is this, on a scale of one to four?
Forbes made the same point about AI agents generally when Jev launched: most of what an agent asks a frontier model to do isn't writing, it's choosing the next step. Until now, every one of those choices has been handed to a model built to write, which then has to compose an answer to say "invoice" or "yes". It works, but it's like hiring a novelist to tick boxes.
A model built to choose
Jev doesn't write at all. You give it something to look at, such as a document, an email or a transaction, and one or more closed questions. It comes back with an answer from the options you gave it, plus how confident it is, in a fraction of a second and at a small fraction of what a conventional model costs. Because it can only answer from your list, it can't invent an account code that isn't in your chart of accounts, or a category that doesn't exist.
TypeSafe calls it a "System One" model, after the psychologist Daniel Kahneman's idea that people have two modes of thinking: fast and intuitive, and slow and deliberate. Jev is built for the fast kind.
Why bookkeeping automation is mostly choosing
Behind the scenes, bookkeeping is almost entirely choosing. What is this document, who is it from, which account does it belong to, is it a duplicate, does it match what was ordered, does a person need to see it? A business with a few thousand transactions a year generates tens of thousands of these small choices.
Some choices are easy and some are hard. Coding a regular supplier's monthly bill is easy. Deciding whether a payment to a director's related company is a genuine business expense takes context, history and reasoning, and sometimes a person. The skill is knowing which is which.
Process AI is built on the idea that every business deserves bookkeeping that checks everything: every line item processed, every supplier's ABN and GST registration verified, every bill matched to its invoice, every payment screened for fraud. No human team could justify that level of checking, which is why even good offshore bookkeeping teams sample and spot-check. AI makes it possible, but thoroughness at that scale takes a lot of processing.
So our platform was designed from day one to use the right model for each job rather than one model for everything. When a new kind of model arrives, the question for us isn't whether to switch to it. It's which jobs it would do better, without giving up anything on accuracy.
What we found in our trial
Our engineering team tested Jev on real Process AI workflows, side by side with the models we use today. Across hundreds of tests, on the tasks Jev completed:
The gap between the top rows and the bottom one is deliberate. We trialled Jev on the quick, high-volume choices: sorting and checking. The harder choices, such as an account code for an unusual cost or whether a related-party payment is a genuine business expense, need more context and reasoning. Those stay with the most capable models we have and with people. We wouldn't hand them to a model built for speed, however cheap it is.
There's also a gap between the headlines and reality. TypeSafe's own figures say Jev can be up to 200 times faster than conventional models. In our real workflows, with everything else a document goes through, it was about three times faster. That's still a real gain, but a very different number.
It's a trial, and we're treating it as one. We're exploring how a model like this could fit into what we do, measured task by task. Anything that wouldn't hold up for our clients doesn't make it in. For us this was never about doing the same work more cheaply. It's about what faster, lighter decisions make possible: more checking, sooner, on more of what passes through a business.
Confidence is the feature, and the risk
Jev's most useful feature is that every answer comes with a confidence score. When it's sure, the work can flow straight on. When it isn't, the item should go to a more capable model or a person. That's how a good junior bookkeeper works.
Confidence scores aren't perfect, though. Independent testers have found that models like this are very reliable when they're completely sure, and much less reliable when they're only fairly sure. They've also found that passing uncertain cases to a bigger model doesn't automatically fix them. So a confidence score is something to measure against real results, not something to take on trust.
The opportunity for small and medium businesses
This is the part we think matters most, and not only for our clients.
Controls that used to belong to big companies are coming within reach
A large company has a finance team, an accounts payable department and an audit trail on every payment. It checks suppliers before it pays them, matches every invoice, and screens for the email asking to change a supplier's bank details. A 20-person business has a bookkeeper a few days a week and a lot of trust. That gap was never about knowing better. It was about cost: checking everything took people, and people are expensive.
The risk isn't hypothetical. Australians reported losing $166.8 million to payment redirection scams in 2025, the second-costliest scam type after investment scams, according to the National Anti-Scam Centre. Among businesses, small businesses reported more scams and higher total losses than medium and large businesses. The Australian Signals Directorate puts the average reported cost of a cybercrime at $56,571 for a small business and $97,166 for a medium one (2024–25). For many SMEs, one missed email is a very bad year.
It reaches well beyond the books
Most of what lands in a small business every day is a choice between a handful of possible answers:
- The inbox. Is this a supplier invoice, a customer query, a quote request or junk? Is it urgent?
- Scam screening. Does this email ask for a payment or a change of bank details? Does it come from the address we normally deal with?
- Customer enquiries. Sales, support or complaint? Who should handle it?
- Compliance basics. Is this a valid tax invoice, with an ABN and the GST shown? Is this supplier's paperwork current?
Each of those is a choice, not a piece of writing, which is exactly what this kind of model is built for.
Five things any SME can take from this
- List your repetitive decisions. Anything a person answers the same way many times a week, from a short list of options, is a candidate. Sorting and checking come first; judgement comes later, if at all.
- Keep a person on the uncertain ones. Use the confidence score as a gate: sure goes through, unsure goes to someone who knows. And don't use a model that can't explain itself for a decision you may have to justify to the ATO, an auditor or a customer.
- Test on your own work, not the brochure. "Up to 200× faster" became about 3× in our workflows. Try any tool on real examples before it touches anything that matters.
- Ask your providers what they use for what. "We use AI" tells you very little. The better question is which AI does which job, and where a person still makes the call.
- Put the gains into doing more. The businesses that benefit most won't be the ones that trim a line on the P&L. They'll be the ones that start checking things they never could before.
That last point is where Jev gets its name. TypeSafe named the model after William Stanley Jevons, the economist who noticed in 1865 that more efficient steam engines made Britain burn more coal, not less, because cheaper power made new uses worthwhile. Forbes reported an early example: a developer had shelved a review step because it cost too much to run on a conventional model, then cleared a queue of more than 9,000 items with Jev for about 32 cents. Work that wasn't worth doing suddenly was.
We think checking is about to follow the same path, and small businesses have the most to gain.
Always testing
We'll keep doing this. Every new model gets tried on real work before it goes anywhere near a client's books, and it earns a place only where it holds our standard. Jev is the latest; it won't be the last.
Where could this work in your business?
The same thinking applies well beyond bookkeeping. Most businesses have processes that run on repetitive decisions: an inbox someone sorts by hand, invoices keyed in twice, approvals chased by email, a spreadsheet that exists only because two systems don't talk to each other.
We help businesses find those processes and put them to work, whether that's automating a workflow from scratch or streamlining one you already have. We'll look at how the work actually flows today, pick out the decisions that could be made faster and checked more thoroughly, and show you where the time and cost go. Every step uses the right model for the job, and a person stays on the calls that need one.
A good place to start is the first step in the list above: bring us the repetitive decisions your team makes every week, and we'll help you work out which ones are worth automating, and which aren't.
Figures are from Process AI's internal trial, based on analysis of hundreds of tests on our own document workflows, September 2026. They cover the tasks the model completed. Results vary with document mix and volume. Scam figures: National Anti-Scam Centre, Targeting Scams report 2025 (published 30 March 2026). Cybercrime cost figures: Australian Signals Directorate, Annual Cyber Threat Report 2024–2025. Developer example: Forbes, 19 September 2026. Jev is a product of TypeSafe AI and is currently in limited early access. Process AI has no commercial relationship with TypeSafe AI.
Where could this work in your business?
Bring us the repetitive decisions your team makes every week. Whether it's automating a workflow from scratch or streamlining one you already have, we'll show you which ones are worth automating, and which aren't, using the right model for each job and keeping a person on the calls that need one. Explore custom AI & automation.
Talk to us about your workflows
