Playing with Jev on SAP BTP
The latest toy is in. Jev, a new type of model developed by TypeSafe, differs from ordinary LLMs in a few ways. The main difference lies in its output: it does not generate conversational text, but it can make decisions. You give it some context, define the available choices and then throw questions at it. You get a structured response with its decision and, crucially, the probability it assigns to that decision.

Not only that, but it also promises to be a lot quicker and tremendously cheap. TypeSafe advertises incredible speed and cost advantages over LLMs for these tasks, although the exact difference depends on the model and workload being compared. Apparently, this is not just introductory pricing: the architecture itself is designed to make these decisions efficiently. (As of writing this, I have made about 120 requests to the API and I have managed to spend 0.2 cents).
I must admit I was not particularly impressed when I first heard about it. Big deal, I thought. The umpteenth LLM being asked pretty please to produce valid JSON from a bunch of Markdown files. But the more I read about it, the more use cases started popping up in my head.
We can now use something that is quick, dirt-cheap and comes with a built-in estimate of how strongly it favours a decision. These probabilities are part of the model’s decision output, rather than numbers it has been prompted to write in a response. That does not make them infallible, but it does give us something useful to evaluate and build thresholds around. And you can feed it unstructured text, JSON or a mixture of both. And it already comes with a Python SDK you can use to duct-tape together a working API?

The thing is, there are many problems that share these three characteristics:
- A human can solve them fairly easily.
- A deterministic process struggles with them.
- An LLM can solve them, but the cost, latency and reliability make you think twice about putting it in production.
I am talking about situations where a decision needs to be made from unstructured information, especially when there is a lot of noise and background context that is difficult to capture in explicit rules. There are countless problems like this in my day-to-day work: checking logs and deciding whether they relate to the issue I need to debug, or getting an error from some system and deciding whether a request should be reprocessed. You get the gist.
This new thing, Jev, as much as it may sound like a middle-aged accountant’s name, promises a middle ground: enough contextual understanding to make a useful decision, with a cost and response time that could make it practical to call from an integration flow. So here we are: let’s try to use it to do something useful.
To retry or not to retry: that is the question
One of the problems I mentioned has bugged me for some time: how do you decide whether a request should be retried or taken out of the retry loop? As a human, it is often fairly easy: retrying the same request with incorrect credentials will not help, while retrying after a timeout might. But this sort of context-informed intuition is hard to apply deterministically when you drift away from APIs with clear, consistent error responses, which are few and far between.
So this is what we want to do: let Jev classify the failure, then use that classification to decide whether the message should be sent back for reprocessing.
Building a simple API in SAP BTP
With the Cloud Foundry environment enabled in an SAP BTP subaccount, you can create a space and deploy applications using different languages and buildpacks. It is flexible enough for this little experiment. So when I read that Jev had a Python SDK, I immediately started planning the use case I just mentioned.
I uploaded my code to this repo. The details are in the README. You can clone it, configure your API key and push it to your own Cloud Foundry space. The instructions for setting up proper authentication are there too.
The interesting bit is that you can adapt this code to make Jev choose in a different context. Here, I tell Jev what it needs to decide and which criteria to use:
TRANSIENT_QUESTION = Noul(
instructions=(
"The state is an error, exception, status code or log entry raised while calling a system. "
"Is this failure transient, meaning the exact same request could succeed if it is simply retried later, "
"without anyone changing credentials, configuration, code or the request payload?"
),
criteria={
"true": (
"Transient: a temporary condition outside the request itself. Examples: HTTP 408, 429, 502, 503, 504; "
"read or connect timeouts; connection reset or refused; temporary DNS failure; service unavailable or "
"under maintenance; throttling or rate limiting; database deadlock or lock timeout; broker temporarily unreachable."
),
"false": (
"Persistent: retrying the same request will keep failing until someone fixes something. Examples: HTTP 400, "
"401, 403, 404, 405, 409, 413, 415, 422; invalid or expired credentials, tokens or certificates; missing "
"authorization or roles; malformed payload; schema, mapping or validation errors; wrong URL or unknown "
"endpoint in configuration; null pointer or other coding bugs."
),
},
)
In this case, I am using a Noul, a yes-or-no question, but you can also ask it to choose between several options. These criteria are a starting point for the experiment; the meaning of an error still depends on the system returning it.
After pushing the app and letting Cloud Foundry go through the initial setup, I did a brief sanity check.
For this input:
{
"error": {
"message": "Server was abducted by aliens. The UFO went away... oh God, it's never coming back..."
}
}
We get this output:
{
"classification": "persistent",
"action": "discard",
"p_persistent": 0.89,
"p_transient": 0.11,
"persistent_threshold": 0.7,
"model": "jev-1.13.0",
"usage": {
"input_tokens": 543,
"output_tokens": 21
}
}
Jev assigns a probability of 0.89 to the failure being persistent. Apparently, waiting for the aliens to return the server is not a promising recovery strategy. Cool!
For something a bit less exciting:
{
"error": {
"message": "Backend timeout after 500 ms."
}
}
We get this:
{
"classification": "transient",
"action": "retry",
"p_persistent": 0.21,
"p_transient": 0.79,
"persistent_threshold": 0.7,
"model": "jev-1.13.0",
"usage": {
"input_tokens": 531,
"output_tokens": 21
}
}
Makes sense!
Time to try something a little closer to a real integration scenario.
Simulating a real(ish) scenario
I came up with a silly REST API that sometimes fails on purpose. Its errors are a mix of very boring-sounding stuff and some exciting, never-before-seen exceptions:
{
"status": "error",
"httpStatus": 409,
"error": "Conflict",
"message": "Two messages tried to update the same record and are now having a knife fight",
"timestamp": "2026-09-28T18:55:56.877646298Z"
}
This will be our untrustworthy destination. Half the time, it works and returns an acknowledgement; the other half, it makes stuff like this up.
To simulate an asynchronous integration, I set up a bridge that receives a message and stores it in a queue:

When the message is read from the queue, it is sent to our evil API, which will sometimes fail. An exception subprocess sends the error information to our BTP app, which returns a classification and the corresponding action.
Let’s see a few examples.
Case one: transient errors
A message is read from the queue and sent to the API, which fails:

The API returns something like this:
{
"status": "error",
"httpStatus": 429,
"error": "Too Many Requests",
"message": "Tenant quota exhausted, try again in: 5 seconds.",
"timestamp": "2026-09-28T19:06:16.398853704Z"
}
We route that to our BTP app and get this decision back:
{
"classification": "transient",
"action": "retry",
"p_persistent": 0.07,
"p_transient": 0.93,
"persistent_threshold": 0.7,
"model": "jev-1.13.0",
"usage": {
"input_tokens": 518,
"output_tokens": 21
}
}
Jev assigns a probability of 0.93 to the failure being transient. That is not a measured 93% chance that the next attempt will succeed, but it supports retrying under our policy. We send the message back to the queue and, in this run, it is processed successfully!

Case two: persistent errors
Suppose the API returns something like this:
{
"status": "error",
"httpStatus": 501,
"error": "Not Implemented",
"message": "Feature scheduled for Q5",
"timestamp": "2026-09-28T20:03:16.027181389Z"
}
Then our app returns:
{
"classification": "persistent",
"action": "discard",
"p_persistent": 0.94,
"p_transient": 0.06,
"persistent_threshold": 0.7,
"model": "jev-1.13.0",
"usage": {
"input_tokens": 519,
"output_tokens": 21
}
}
The message is taken out of the retry loop. In this demo, the action is called discard, but a real implementation could route it to a dead-letter queue (DLQ) for investigation and later recovery. Q5 might take a while.
What’s next?
This is a working proof of concept, but a few convincing examples are not enough to call it production-ready. The next step is to find out how it behaves when the errors stop being funny and start looking like the contradictory, incomplete messages we actually get from real systems.
I would start with a collection of real failures, labelled according to whether retrying was useful, and compare Jev with a simple rules-based baseline. The interesting cases are the ones where the status code do not show the full picture. I would also check whether its probability scores hold up on that data before treating a threshold such as 0.7 as anything more than a convenient starting point, which is what I did.
Then there is the surrounding retry policy. A useful classification does not tell us how long to wait, how many times to retry or whether repeating the request could create a duplicate. Those decisions still need explicit rules. Uncertain cases should have somewhere to go, and a persistent failure should usually mean “keep this for investigation”, not “go ahead and delete a business message”.
There are also a few practical questions to answer:
- Performance-wise, what are the end-to-end latency and throughput under realistic load?
- What happens if the classification service itself times out or becomes unavailable?
- which parts of an error or payload can we send to a third party?
- How will pricing, model updates and changes in behaviour affect the integration over time?
Local inference is where I think these new systems will find more practical application. Laya, for example, offers a locally runnable approach to typed decisions, and it’s so small it runs in a browser! That could give us more control over deployment, data and model versions, although we would also take on the work of hosting, evaluating and maintaining it. I would want to compare it with Jev on the same errors before drawing conclusions about either.
What makes this interesting to me is how closely it fits a recurring integration problem: we often have enough information to make a decision, but expressing that decision as a maintainable set of rules is awkward. A small, fast model could be useful at precisely that point, provided we measure its mistakes and design the process to handle them.
For now, Jev has passed the “can I build something useful with this?” test. Next comes the harder one: can it make better retry decisions than the rules I would otherwise maintain, at a cost and level of reliability I can justify? That is an experiment worth running.