In an earlier tweet, I wrote that Jev mattered more than I had initially thought. Part of that comes from the way its interface encourages us to separate decisions from generation. After evaluating this approach, we found that the separation itself could help in some cases, even without introducing a new model. In Nowledge Mem’s memory review step, for example, we saw better results just from restructuring the task this way.
I’m also excited about Jev-like models as a kind of intelligent coprocessor: “System 1” models capable of generalization and in-context learning. Many tasks require intelligence without requiring generated text. A model may need to understand a whole conversation, then make just a handful of decisions. If different subsystems can handle those decisions, we have a way to offload work from the “brain,” much as learned abilities can become the responsibility of reflexes, the cerebellum, or the autonomic nervous system.
Then there is the more immediate value of Jev as a low-latency, prompt-based classifier. It can potentially improve judgments currently handled by heuristics, or bring semantic understanding to parts of a system where latency and cost have made it impractical.
These ideas ran through a little over a week of experiments. We’ve now released the version of Nowledge Mem with Jev support, so we can discuss what we observed alongside what we actually put into use.
Semantic decisions inside a memory system
Nowledge Mem is the context layer infrastructure we’re building, with both software and services. Mem helps people and teams preserve and organize context from their work with AI, so different agents can use the knowledge they’ve accumulated. Extracting memories from conversations, organizing saved knowledge, and selecting relevant context during retrieval all involve decisions that depend on understanding meaning.
“This command succeeded” and “Run this check before every deployment” are both new pieces of information. One records the outcome of an operation; the other may be a working agreement that should still apply in the future. Length, keywords, and timestamps tell us little about that distinction or whether the information will remain useful.
Before saving a memory, we also need to check whether its source supports it, whether it goes beyond what was actually said, whether we’ve saved it already, and whether it updates an earlier decision. Those judgments shape the context agents will encounter later. An incorrect memory can be retrieved and cited repeatedly, and may even become the basis for further conclusions. We need these checks in many places, but calling an LLM for each one adds cost and waiting time throughout the system.
Jev is a decision model from the TypeSafe team. A request provides some material as state and a set of questions; the model returns answers within specified bounds, along with their probabilities. A question can ask whether something is true, ask the model to choose from a fixed set of options, or request a score against a defined criterion.
That interface lets us define the decisions in our memory system independently, ask several questions about the same material in one call, and have different models answer them. We can organize and test the criteria separately. Work that needs generated text—memory content, summaries, entity names—stays with the LLM.
Separating decisions from generation, with the same model
Inspired by Jev, we started with a set of experiments on candidate memory review. After an LLM extracts information from a conversation, we check whether it is worth keeping long term, whether the source supports it, and whether the model has overinterpreted anything. Our existing implementation already required these checks, but the model did several jobs in one call: decide whether to keep a candidate, revise its title and content if necessary, and return the resulting memory as JSON. To prepare for Jev, we turned the judgments into individual questions, each with explicit criteria and allowed answers, and asked the model to answer only those questions.
We put this interface into DecisionClient, so Jev and chat models could answer the same set of questions. That also let us keep the model unchanged at first and compare how we organized the task, to understand where any gains came from.
We had the same Kimi k3 model review 40 candidate memories in two ways: the original review process, and a version that returned only answers to the question list. Both made one call per candidate:
The model improved at judging long-term value and identifying overinterpretation. Memory type accuracy stayed the same. An earlier comparison had shown a large improvement in type classification as well, but it did not repeat in this run, so we cannot treat it as a reliable gain. Type accuracy was measured only on the 19 candidates labeled as worth retaining long term.
Those perfect scores apply only to these 40 examples. We also clarified some criteria while turning the review into separate questions, so this comparison alone cannot tell us whether the improvement came from the task format or clearer instructions. To examine that, we compared five variants in the same ablation run, including one that asked for a reason with every answer:
| Variant | Overinterpretation F1 | Long-term retention F1 | Relative cost |
|---|---|---|---|
| Original review | 0.27 | 0.83 | 1.10 |
| Original format, with commonly confused cases made explicit | 0.42 | 0.86 | 1.19 |
| Question list, with a one-sentence reason per answer | 1.00 | 1.00 | 1.41 |
| Question list, answers only | 1.00 | 1.00 | 1.00 |
| Question list, one call per question | 0.82 | 1.00 | 2.83 |
The last column compares the cost of a complete review of one candidate, calculated using Kimi’s API prices at the time. The answers-only question list is the reference, with a cost of 1.
Making the confusing cases explicit in the original prompt already helped. But it did not match the version that specified each judgment, its criteria, and its answer choices separately. Once we had that question list, asking the model to explain each answer increased cost without improving the review scores.
We ran the same comparison for write-quality checks and standalone memory type classification. The results were similar: requiring reasons did not improve the quality-check scores or type classification. Across the three tasks, it increased total cost by about 35%, and median latency per call roughly doubled:
In these parts of the system, the program uses the model’s decisions. Writing reasons added cost and latency, and these comparisons showed no corresponding quality gain. That gives us a reason to reconsider which text we actually need the model to generate. The experiment was about whether writing reasons helped; a model that does not output reasons may still reason internally.
Once each decision is defined as a question, the way we group calls matters too. Reviewing one candidate requires six questions about the same material. Asking them in six sequential calls took about 50 seconds at p50 for the complete review. Answering all six together took 11.6 seconds. Amortized across the questions, that is a reduction from 8.4 to 1.9 seconds each, with about 65% lower cost:
We had not changed the chat model in any of these experiments. We were comparing task formulation, whether to generate explanations, and how to group calls. Building an abstraction for Jev gave us a way to evaluate those choices independently and find some gains there. That was the first observation I wanted to share in the tweet.
Which generation work can we avoid?
Alongside evaluating the decisions themselves, we took inventory of the model calls across our background processes. We estimated a heavy user’s monthly costs at 20 conversation threads and 60 memories per day, to understand how much of the overall spend a decision model could affect.
In that scenario, decision tasks accounted for more than half the calls but only about a quarter of the cost. Most of the spending went to generation: extracting memories from conversations, extracting graph information, and backfilling data. Some of that work began before we had established whether it was necessary:
Replacing the model that makes decisions therefore offers limited direct savings. But a model like Jev, which can understand context at relatively low cost and latency, makes it more practical to check whether expensive generation is needed before starting it. When a conversation changes, for example, we can first ask whether the new content contains knowledge worth keeping that we have not already saved. We then decide whether to ask the LLM to extract it.
Using the same workload assumptions, take the original process as a cost of 100. If we keep the existing model, move decisions ahead of generation, and batch work that can be done together, the estimated cost falls to about 37. Assigning those decisions to Jev brings it to about 26, using the Jev team’s prices at the time:
Under these assumptions, about 85% of the savings come from changes to the workflow, including batching unrelated to Jev. The other 15% comes from the difference in model prices. The estimate depends heavily on usage: how much content can actually be skipped, and how much work can be grouped together, need to be measured with real traffic.
In small integration tests against the actual services, we observed roughly 10%–19% lower cost per conversation thread, well below the modeled reduction. In those tests, the original model still handled candidate review and we only recorded the decision model’s answers at that stage. Several other stages used its answers to determine the work that followed. We have not yet had enough time to run an A/B comparison of production users’ full monthly bills.
What happens after a decision matters here. If most content still needs extraction, or if a chat model must double-check the decision, the new call can simply add overhead. We also need to revisit skipped content and check that it did not contain knowledge we should have saved.
What a decision changes about saving a memory
Screening content before extraction requires two judgments: the new information is worth keeping long term, and it has not already been saved. If the new content only reports a transient operation result or a temporary status, or repeats knowledge we have already stored, it does not need to go through LLM extraction.
The material we use to check for existing knowledge is critical. A decision appearing in a summary does not mean it has been written to the memory store. If the previous write failed and we compare against that summary again, the model may correctly understand the content and still skip knowledge we have never saved. The comparison must use memories that were successfully written.
We also need to distinguish “could not make a decision” from “no new knowledge.” If the decision call times out or fails, we continue through the original extraction process. A missing answer cannot be a reason to skip the content.
In candidate review after extraction, a candidate accepted by mistake can still be caught by deduplication or write-quality checks. A candidate rejected by mistake never reaches the save process. We need different conditions for accepting the model’s answers depending on those consequences. One confidence threshold for every judgment will not do.
We tried a conservative approach: use the model’s answer only when it is sufficiently confident that a candidate should pass, and send rejections or uncertain answers back through the original review. Across three small comparison runs, we saw no mistaken acceptances, but the amount of correct knowledge ultimately saved was sometimes higher and sometimes lower than with the original process. Cost did not fall either. Candidate review therefore still runs in shadow mode: we record the decision model’s answers, while the original process makes the actual review decisions.
When we use decisions to skip generation, the final check is whether the knowledge that should have been saved is still there. We have to connect what the model saw, which answers the program acted on, and what was finally written. Only then can we tell whether one fewer call removed unnecessary work or lost a memory.
Memory retrieval
Memory retrieval was one of the first places we thought Jev could help. We’ve written about Mem’s search pipeline before. The design raises several questions that benefit from understanding a query: how to weight the different retrieval paths, whether HyDE would be useful, and whether the query contains time constraints. Latency makes it hard to give all those decisions to a chat model in the default fast search path.
We also tested whether input caching could reduce that wait. We sent six consecutive decision requests with an identical prefix to the same reasoning chat model. About 88% of the input was cached for the last five calls, but each model call still took 6 to 33 seconds. These were isolated model calls in an experiment, not Mem’s retrieval latency. Default fast search does not include this step.
Our Jev integration started with understanding time information in queries. As we discussed in our article on time in memory, when something happened and when it was recorded are different things. Phrases such as “last year” or “after the last migration” also need to be interpreted in context.
The existing chat model still handles general intent analysis; the decision model adds temporal interpretation. In Memory Deep Search, if a decision service is configured, we also try using it for reranking when it identifies temporal intent and returns a usable result within the time budget. Otherwise, we use the existing reranker.
For fast search, we tried starting retrieval and temporal intent analysis together, allowing a bounded wait for the analysis before ranking and proceeding with the original retrieval results if the analysis exceeded that budget. In our tests, retrieval finished sooner, and the decision model’s answer did not arrive in time to help rank the results. Decision models are therefore off by default for fast search. For now, this integration is in deep search.
Even when we have a temporal interpretation, a model-inferred range only influences ranking and is returned as a suggested filter. Only filters the user explicitly provides are enforced. An incorrectly inferred time range could otherwise remove evidence from the results, leaving the downstream agent with no opportunity to notice what went wrong.
Can local models make these decisions?
Work like Laya is particularly exciting to me. If a model running in the background on the user’s computer could screen content before extraction or interpret search intent, those tasks could happen locally without waiting for a remote call each time. We evaluated several community open-weight models to see whether their capabilities, latency, and resource use would fit these jobs in Mem.
We first tested three encoder models—Laya, Von, and Verdict—with a fixed set of questions and the same Chinese and English examples. Each model completed 290 calls on an 8-core CPU machine with no GPU.
We set the requirement in advance: on each task, both Chinese and English scores had to be no more than 0.05 below the chat baseline, measured by accuracy or F1 as appropriate. None of the three models met that requirement on any of the seven tasks. Their p50 call latency ranged from 0.2 to 3.3 seconds, and peak process memory was 1.7 to 4.1 GB. They were also some distance from what we would want for a model that stays running on a user’s computer.
Our friend PsiACE, one of Bub’s authors, is working on Dohnuts, another interesting approach. It uses Qwen3.5-0.8B with LoRA and a dedicated scoring head to learn decision tasks, keeping the base model’s weights unchanged during training.
On the same questions, Dohnuts beat the best of the three encoders on five of the seven tasks. For write-quality judgment, the best encoder scored 0.17 F1 and Dohnuts reached 0.74. For memory type classification, the best encoder achieved 0.48 accuracy and Dohnuts reached 0.70.
Both tasks involve concepts specific to Mem, and Dohnuts did substantially better on them. We did not see similar gains on tasks that require reading a full passage and comparing it with other content, and it still did not meet our overall requirements for integration. Separate experiments would be needed to tell how much of the improvement comes from the Qwen base model, the training data, or the training method.
We then narrowed the scope to six apparently simpler tasks, such as deciding whether a tag can be reused, whether a web result is relevant, or whether a page lacks substantive content. Each task had 40 examples in total, covering Chinese, English, and Japanese. Even here, no local model met the earlier requirement in both Chinese and English on any task. The closest was Dohnuts on web relevance: its Chinese F1 was 0.94, against a chat baseline of 1.00.
Some of these tasks end in a yes-or-no answer, but arriving at that answer still requires understanding the content and its context. That is part of why I want to see whether larger base models and different training data can make these models useful for a wider range of decisions.
Using models to improve heuristics
We also ran common rules—word lists, minimum text lengths, and writing-system detection—through the same six tests. The rules fell behind the chat model on all six tasks, and behind the smallest local model on four. Judging whether a page has substantive content by its length, for example, also excludes some pages that are short but useful.
Permissions, budgets, date arithmetic, and bounds checking should of course remain in code. But deciding whether a page contains something useful requires understanding it; page length is only a proxy. Testing rules and models on the same examples lets us compare where each makes mistakes.
The baseline therefore depends on what the model would replace. Replacing an existing chat call means meeting its quality requirements. Improving a heuristic means reducing the rule’s errors at an acceptable cost. The local models we tested have not met our integration requirements, but the comparison with rules gives us a reason to keep exploring.
What is running today
You can now configure a decision service under “Fast decisions” in Settings. If you bring your own API key, you can add TypeSafe AI. For managed AI accounts, the service supplies the configuration. Once enabled, decision models are used at the following stages, with some differences between desktop and Cloud:
| Stage | Current use |
|---|---|
| Decide whether a conversation is worth extracting memories from | Decision models are in use on desktop and Cloud |
| Check whether new conversation content contains unsaved knowledge worth retaining long term | Desktop uses the decision to determine whether to proceed with extraction |
| Deduplication, write-quality checks, and type classification | Desktop uses decision models for these judgments |
| Review candidate memories for long-term value and source support | Desktop records decision-model results; the original review process still decides |
| Temporal interpretation and reranking | Desktop deep search uses temporal interpretation and attempts decision-model reranking when a temporal intent is identified and the result is usable; off by default in fast search |
| Screening before graph extraction | Cloud uses a decision model for screening; the LLM still extracts entities and relationships |
Beyond replacing existing model calls, we are also testing decision models for recognizing content injected by the host tool, deciding whether a tag can be reused or a page lacks substantive content, and identifying the language of a description. These judgments are currently recorded only; they do not affect subsequent processing. We will continue comparing them with the existing approaches. If no decision service is configured, these additional calls do not happen.
Whether the memories we ultimately save are more accurate, users’ full monthly bills are lower, or search is faster still requires production A/B tests covering the complete workflows. We do not have those results yet.
If you are doing similar work, I would strongly recommend keeping the original implementation as a baseline. Test the separated decisions with the same model first, then keep the questions fixed while changing models, and finally evaluate what happens across the complete workflow after integration. That lets you distinguish gains from reorganizing the task from gains due to the model, and check whether either improves the final outcome. If you are trying to improve a heuristic, include the current rules in that comparison too.
Where I’d like to go next
We recently came across the Jev-Mem paper from UT Dallas. They also use decision models in memory construction and retrieval, leaving complex reasoning and answer generation to an LLM, which is close to what we have been working on. They evaluate the final question-answering results on LoCoMo. Our work here has focused more on how to formulate each decision and whether the system can act directly on the answer. I think the two approaches complement each other well.
This work has made my interest in “neural subsystems” more concrete. Once we separate decisions from generation, we can ask what level of model capability each decision really needs, whether it can run locally, when the cloud is a better fit, and how soon the rest of the system needs its answer. If these System 1 models can continually interpret and assess incoming context at low enough cost, we could do work that has been hard to perform frequently.
We will keep testing local Jev-like models, and try decision models for real-time channels, Feed cards, and cross-language tags. I am also curious about more personal capabilities, such as learning what is worth preserving for a particular person. I would like to know how far dedicated training could take that. So far, our work has been on question design and threshold calibration; we have not trained such a model.
I also do not think every coprocessor model should be as small and fast as possible. As I mentioned in the tweet, Jev-like models with larger-scale pretraining and stronger generalization could have many uses beyond simple reflexes, even at higher latency. The decisions in Mem that require understanding a whole context make me particularly interested in that direction. I also want to explore whether decision-oriented post-training of open-weight LLMs can produce those capabilities, or whether it calls for other training methods.
Robotics and embodied intelligence are another direction I find interesting. I think a purpose-built harness combining several different Jev-like models with a frontier LLM, including at least one model that understands visual input, could let these models participate in a robot’s feedback and decision loops. I do not know the interfaces between robotics subsystems well. But my guess is that such a combination could make it easier to develop capabilities that previously required expensive training.
That raises another research question: how can an LLM and its harness coordinate several subsidiary models with low latency? What differences in capability and latency would we see if we designed those models together as a system, compared with connecting them programmatically through a harness? An imperfect analogy is an end-to-end voice model versus a voice agent framework assembled from separate components.
Returning to the nervous-system analogy, AI seems to have arrived at “brain-level” models before exploring general-purpose neural subsystems. I find that order fascinating. I would love to discuss these ideas with others working on memory, agents, and models, and to see more work exploring them.
Experiment notes
The chat ablations used Kimi k3, with 40 examples per task. The charts show results from the same evaluation run; small differences need repeated experiments before we can draw conclusions from them.
Chat-model costs were calculated using Kimi’s API prices at the time, so we could compare approaches on a common price basis. We actually ran the experiments through a subscription, so these figures are not our paid bills.
At the Jev team’s request, we cannot disclose Jev’s own performance figures. The model evaluation numbers in this article therefore come from chat models and open-weight models.