늘모자란, 개발

Do Fast Decision Models Actually Make AI Agents Faster?

Testing KEV, JEV, LAYA, and PPLX—and revising the decision integration through v12

Contents

What if a large model handled long articles and code, while a faster model handled the small decisions along the way? Those decisions include choosing the next file to read, finding the right skills, and checking whether the available evidence is enough to finish a task.

That question led me to test KEV 0.8B, 4B, and 8B, the JEV API, LAYA, and PPLX Decider 27B in turn. I compared modified versions of a real operations skill and connected decision models across the agent workflow. Along the way, I found more problems in inputs, hooks, validation, and recordkeeping than in model performance alone. The integration client evolved through v12.

The fastest model was not necessarily the most useful. A larger model getting every answer right in a small test did not establish near-100% accuracy in real work. Even when a decision API responded in around 100ms, the full task could take longer. Conversely, a model that took several seconds could reduce waiting time when several questions reused the same observations.

The current experimental setup calls KEV 0.8B and PPLX 27B in parallel with the same decision packet. It compares their answers without automatically treating either as correct. In a short installation check, the first call took 2.829 seconds, and a cache hit with the same state took 1.141 seconds. Questions with fresh state from actual work generally took longer.

This is an experiment log, rather than a success story about a system designed correctly from the start. It records what I tested, which numbers were easy to misread, and how I separated and repaired failures. The measurements and implementation status are based on records checked on October 3, 2026.

1. Separating the model that writes from the model that decides

Agents make short choices throughout a task. After checking a server, an agent might decide whether to read more logs, look for an existing automation first, use test results as evidence for deployment, or ask for clarification because the evidence is incomplete. A single choice can sometimes prompt a long stretch of reasoning by the primary model.

The goal was to narrow the alternatives enough to choose the next action quickly, rather than replace document writing or debugging with a small model. “Fix this system” is too broad. “Given this error and file state, should we first inspect a directory conflict or the network?” can be a bounded question.

KEV, TypeSafe's JEV, LAYA, and PPLX Decider were the candidates. I was interested in an approach that accepts a natural-language state and question and returns a constrained decision. Their implementations and supported capabilities differ, so similar interfaces do not guarantee equal performance or equivalent probability semantics.

Most of the experiment used choice questions. An input contained the observed state, a decision question, and alternatives describing actual actions. The response provided a choice and a distribution. This was a different task from asking a general chat model for a long explanation and parsing its final sentence.

The initial hypotheses were modest:

To test these hypotheses, I thought existing work would be more useful than a large collection of examples such as shopping recommendations. The first target was an operations skill for checking LAN hosts and services.

2. Comparing modified versions of a real operations skill

The first experiment kept the existing skill separate from candidate versions using decision models. Each performed read-only checks of services and processes, then evaluated five claims against the evidence collected.

Claim to checkObservation requiredReference judgment at the time
Is the model API process running?Observation of that processSupported
Does the web UI or model proxy return an HTTP response?HTTP result from that addressSupported
Did an inference request complete?An actual inference request and its returned resultSupported
Is the specified agent runtime running?Process result for the target runtimeSupported
Has message delivery also been confirmed?Evidence confirming both sending and receiptNot confirmed

The last claim mattered. A running process does not establish message delivery. This test sent no message and had no delivery receipt. An HTTP response from a web service likewise did not mean every feature worked in the browser.

I observed full runs of the existing skill and versions using KEV 0.8B, KEV 4B, and JEV.

Skill configurationFull execution timeDecision model's initial judgment in that runFive claims in the final answer
Existing skill167.070 secondsNot applicable5/5 matched
KEV 0.8B candidate174.990 seconds4/5 matched5/5 matched; one item corrected
KEV 4B candidate181.871 seconds5/5 matched on newly collected evidence5/5 matched
JEV candidate188.015 seconds5/5 matched on newly collected evidence5/5 matched

Each configuration ran once. The material read, tool strategy, runtime state, and waiting time differed. The table shows what runs with decision models looked like, but it is not a controlled comparison establishing a percentage reduction in full task time for any model.

In the observed runs, KEV 0.8B took 7.920 seconds longer than the existing skill. Adding a decision model did not automatically make the task faster. The existing run made 11 shell calls, while the 0.8B candidate made 15. Preparing decision inputs and checking results added work. Review by the primary model was also needed to preserve the final answer's quality.

The first 0.8B decision request took 2.433 seconds. For the final claim, which lacked delivery confirmation, the model returned a value of 0.5127 for true. Under the simple 0.5 boundary used at the time, that value could be classified as “delivered.” The primary model checked it against the original observations and corrected it.

Although 0.5127 does not look particularly certain, turning that number into a binary choice makes it an actual conclusion. The boundary was not a policy calibrated for this task. This was the first example showing the gap between receiving a score and verifying a decision that was safe to act on.

Fixed observation sets produced different results

The full skill runs collected fresh evidence each time. To compare the model judgments more directly, I froze two observation sets from actual checks and submitted them again.

ModelFive claims in observation set AFive claims in observation set B
KEV 0.8B4/5 matched3/5 matched
KEV 4B3/5 matched5/5 matched
KEV 8B4/5 matched3/5 matched
JEV jev-1.13.04/5 matched4/5 matched

Both sets came from checks of similar services. Five claims in each set did not amount to ten randomly sampled, independent problems. Even so, the results did not support the expectation that 4B or 8B would necessarily outperform 0.8B on this small task.

Nor did they contradict the earlier 5/5 results for 4B and JEV. Fresh evidence collected in a new skill run and a frozen observation set are different inputs. Ignoring that distinction would make the test results look like consistent model performance that they did not establish.

3. Why I kept 0.8B after testing 0.8B, 4B, 8B, and JEV

I did not choose 0.8B because it had been the most accurate from the start. It was quick to call locally, used little memory, and the larger models did not show a consistent improvement in these tests.

Repeating the same observations took about 0.331~0.350 seconds on 0.8B. A fresh observation set took 2.888 seconds. The difference between repeated and fresh inputs was substantial even early on. Faster repeated responses alone did not establish which internal cache contributed, or by how much.

On a Mac with fewer available resources, 4B hit the 30-second limit twice. A separate inference observation took 81.01 seconds. I observed model computation waiting on the Metal side alongside substantial swap use, but that did not prove the sole cause of the timeouts. Some attempts also failed while constructing the input, before a decision request was made. I kept those preparation failures separate from cases in which a model returned an incorrect answer.

4B ran on another Mac with 32GB of memory. In the BF16 configuration, the first fixed input took 10.625 seconds and repeats took 1.035~1.094 seconds. A skill decision with fresh evidence took 11.271 seconds. Being able to run a larger model and finding it useful enough to adopt were separate questions.

I also tested 8B on a 32GB Mac. The first observation set took 22.659 seconds and the second took 25.306 seconds. Repeats took about 1.9 seconds and 2.0~2.2 seconds, respectively. All six requests completed, but agreement with the fixed references did not improve over 0.8B.

The 8B model was a separate checkpoint available at the time. It was not simply a larger parameter count on the same base model and execution path as 0.8B and 4B. The 8B test used Torch/MPS BF16, a different backend from the smaller-model tests. The model list in today's KEV repository should not be projected backward onto that historical setup.

The JEV API was very fast. Across six requests repeating the two fixed evidence sets, the median was 272.0ms and the range was 251.1~305.2ms. Those times included the client and communication; server-side caching was unknown.

However, it repeatedly failed to recognize inference requests that had actually completed in both evidence sets. JEV matched all five claims in a fresh skill run, but matched 4/5 in each fixed-input comparison. A fast, well-organized paid API did not guarantee correctness on a particular task.

I did not inspect actual billing records in this experiment, so I cannot calculate JEV's cost efficiency numerically. Local KEV does not cost 0 simply because there is no API fee, either. It consumes hardware, power, resident memory, and setup and maintenance time.

There was still a reason to keep 0.8B as the default candidate for small, repeated decisions. The choice was closer to “keep a lightweight default and validate it in actual use” than “its accuracy has been adequately proven.”

4. Expanding from one skill to decisions across the agent

After the first skill test, the goal grew. I wanted agents to consult a local decision model whenever they made an important choice, rather than only within a particular validation skill.

The intended uses included candidate selection, material filtering, classification and prioritization, next actions, deciding whether to investigate further, and simple completion checks. The primary model continued to handle writing, coding, complex problem solving, and assembling final deliverables.

Initially, I added rules to an existing validation skill. But tying a requirement for decisions across the whole agent to one validation skill did not fit the scope. I created a separate decision skill, connected it to global and default instructions, and added Codex and Hermes lifecycle hooks.

Hooks run at events exposed by the runtime, such as prompt submission, before and after tool use, or before model calls. Those events do not represent every agent judgment. Ordinary configuration hooks also cannot fully intercept each moment of internal reasoning.

Global instructions, skills, and hooks therefore could not guarantee that every decision passed through the model. They could remind the agent to consult it where possible and record actual consultations. “Installed,” “passed trust settings,” “applied by the currently open GUI session,” and “called for an actual decision” were separate checks.

This distinction also mattered in later failure analysis. A successful hook trust check in a fresh CLI process did not verify that an already-open app session had applied the settings.

The first monitoring window showed 92.7% insufficient evidence

Early Windows observations, with substantial installation and test activity mixed in, contained 561 calls. There were 559 successful consultations, of which 518 returned insufficient_evidence: 92.7% using successful consultations as the denominator. Two calls were recorded separately as blocked.

Median response time was 675ms and p95 was 1,286.7ms. Calls returned fairly quickly. But when almost every answer reported insufficient evidence, it was difficult to call the setup useful for choosing the next action.

I did not call 92.7% the model's error rate. Inputs might have lacked necessary evidence, and abstaining on difficult questions might have been appropriate. Still, attaching calls had plainly not accomplished the goal by itself.

At the time, hooks mostly asked generic next-action questions based on lifecycle events. They did not adequately convey decisive observations and concrete alternatives from the actual task. The model was not being given a well-formed problem it could judge.

5. More context did not solve the problem by itself

In v2, I passed more task context, including the current request and visible tool results. I expected the previously missing history to make the question answerable.

I replayed three real past requests. Event-only inputs produced abstentions on all three. Candidates with history also abstained on all three, as did a version that changed how the history was packaged.

Event inputs took about 460~470ms, while candidates with history took about 1.2~1.4 seconds. More context increased waiting time without improving the decisions. Because the replay also changed the alternatives, this was not a controlled experiment isolating context length.

Inputs that isolated a specific decision within the same task behaved differently. The important part was preserving which facts distinguished which alternatives, rather than sending more of the entire conversation.

For example, “What should happen after tool use?” provides little basis for choosing investigation, repair, or completion. If the observations establish a file-versus-directory conflict, an already-scheduled retry, and a need to resolve the conflict first, the questions can be separated. “Is a retry missing?” and “Should the conflict be fixed?” can be evaluated independently.

That experience changed the hooks' role in v3. They became short checkpoints reminding the agent to make a concrete consultation when a decision was needed. I stopped sending generic model questions based only on lifecycle events.

An actual consultation occurs when the primary model gathers visible evidence and builds an explicit decision packet. I also separated the counts so checkpoints without model calls would not be included as consultations or abstentions.

6. What changed from v1 through v12

The v1~v12 labels in this article refer to the integration client and instructions for requesting decisions and recording results, not model training versions. The early setup did not have version labels organized as precisely as the current one. The table groups the surviving deployment and repair records by function, starting with the initial integration.

StageProblem foundChange madeWhat the change still did not prove
Initial integrationRepeated generic questions at lifecycle eventsLocal model consultations connected to global instructions, skills, and hooksWhether every actual decision was consulted
v2Task context was missing or passed in wrappersVisible requests and tool history, with context-version recordsBetter decisions from more context alone
v3Checkpoints and actual consultations were mixed togetherHooks as reminders; concrete packets as separate consultationsObservation of every internal judgment
v4Failure stages and follow-up causes were unclearDiagnostics by stage, readable packets, and follow-up recordsIndependent correctness validation from records alone
v5Question lists and input usage did not alignExplicit-ID list support and revised input validationSemantic correctness from syntax compatibility
v6Follow-up file input and the actual location of alternatives were used incorrectlyCorrected follow-up input; rejected unsupported question fieldsElimination of order-dependent decisions
v7Validation and communication failures were grouped togetherValidation-only runs separated from consultations; communication failure subtypesRoot causes from error types alone
v8Requests continued after file creation failedDistinct file errors; instructions to submit only after successful preparation and validationEnforcement of dependent execution across the whole runtime
v9Adoption and tool success were counted as correctnessDecision IDs linked to action and requirement references; separate outcome countsComparative advantage of a successful action
v10Explicitly forbidden alternatives remained available to the modelSource-backed forbidden alternatives removed before inferenceImproved model understanding of constraints
v11Original inputs from past failures were difficult to reviewBounded retention of reviewed requests, answers, and follow-ups on their owning runtimesAccuracy from the existence of case files alone
v12A wish to compare the small model with a stronger oneThe same packet submitted to KEV and PPLX in parallel; separate answers, latency, and disagreementsCorrectness or task savings from model agreement

Rather than designing all these features at once, I repaired problems as they appeared in actual calls and user-visible screens. Many changes improved how failures were interpreted and recorded, rather than directly improving decision quality.

Does correcting the request format count as an improvement?

Putting alternative descriptions in an options field the model does not read, or sending a follow-up review value as an object instead of a string, makes a request invalid. Distinguishing errors more precisely and preventing them in preflight is an improvement: users can find the cause sooner, and failed preparation is less likely to lead to unnecessary model calls.

But those repairs do not establish better semantic judgment by the model. A model reasoning better and finally receiving the required input correctly are separate things.

The two historical invalid_arguments cases checked in v6 were problems with commands recording follow-up outcomes, rather than inference failures. Including them in a model error rate would produce a meaningless number. In v7, I excluded validation-only runs from consultation and failure counts.

v8 distinguished missing inputs, paths that were directories, access-denied errors, and read failures. The trigger was a case in which the program creating the input file failed, yet the next command still attempted submission. File preparation, local validation, and network submission needed to proceed only after the preceding stage succeeded. The instructions require this, but they do not forcibly intercept every shell command the agent executes.

A poorly written question can trouble a large model, too

An ID such as retry_status does not help if the actual question is ambiguous. A question must make sense without its ID. Each alternative must describe the action it represents.

Names such as “an incorrect repair that ignores the current state” and “the recommended safe repair” steer the model toward a conclusion. It is better to describe the actions neutrally and place requirements and observations in the state.

The same applied to skill selection. Instead of asking the model to pick from names alone, the agent had to read actual candidate descriptions and relevant instructions. Multiple skills can apply to one task, so independently useful workflows should not be forced into a single-choice question. No applicable skill was also an allowed result.

There were limits to continually adding instructions. On one runtime, default instructions grew to 21,020 characters, exceeding the 20,000-character limit, and normal validation rolled the change back. I removed repeated explanations across versions, reduced the text to 7,087 characters, and validated it again. This repair made the instructions fit the runtime's allowed size; it was not a measured token-saving result.

While reviewing this direction, I consulted TypeSafe's official skill. I learned from its approach to composing structured decisions in code, but did not import its example JEV thresholds as if they were validated local KEV policies.

I also reviewed omo-jev-plugin. Rather than add the entire external plugin, I considered how to connect decisions to actual actions and follow-up outcomes within the existing system. Its confidence criteria and retry rules had not been validated in our tests, so I did not adopt them unchanged.

7. Hook errors that looked like model failures

During operation, hook exited with code 127 appeared repeatedly on screen. Direct inspection showed that a separate OMX hook pointed to an old Node executable path that no longer existed. After a Node upgrade, its version-directory path had gone stale. A harmless version-check command reproduced the same 127 exit code.

This was not KEV making a wrong decision or refusing to answer. The command needed to run the hook could not be found. As requested, I removed the unnecessary OMX integration and checked the result. The screen alone, however, did not establish that every individual failure row came from that same hook.

At another point, PreCompact and PostCompact hooks produced invalid JSON output errors. A separate recording hook returned a context-output format that those events did not accept. I changed it to preserve successful recording while leaving stdout empty. Isolated tests of the configured commands passed, but I did not observe the next natural compaction event in an already-open conversation.

These issues belong in a decision-model experiment because hook execution, trust settings, input validation, communication, inference, and follow-up recording can all appear as similar failures to someone using the integration. Without separating the layers, changing the model leaves the original problem in place.

Failure layerExampleWhat to check
Execution commandStale executable path; exit code 127Actual file and execution result of the configured command
Hook outputJSON that does not fit the eventOutput contract for that event
Trust and applicationRegistered, but execution permission or application to the current session is unclearSeparate checks of a fresh process and the current session
Local preparationMissing file, failed JSON creation, invalid fieldExit status of production and validation stages
CommunicationTimeout, connection refused, connection reset, name-resolution failureRequest stage at which it occurred
Model judgmentA choice conflicting with evidence; an answer that changes with orderThe same original packet and actual requirements
Result recordingFailed follow-up command; no decision ID to linkOriginal consultation ID and recording success

More detailed communication error names do not solve the cause. connection_refused identifies a failure type. Whether the service was down, the address was wrong, or another prerequisite was missing still needs investigation. I did not adopt longer timeouts and automatic retries as a response to an unknown cause.

8. Changing how constraints and follow-up outcomes were recorded

A small model could choose an action that violated clearly stated requirements. In v10, I added a way to exclude actions explicitly forbidden by the user before asking the model.

For example, if the requirement is to preserve an existing roster and an alternative replaces it immediately, that alternative can be excluded by citing the requirement. The uncertainty alternative remains. If no concrete alternative survives the exclusions, preflight stops.

This does not improve the model's ability to understand constraints. It is caller-side enforcement of explicit conditions. Nor should competing alternatives be removed merely to get a preferred answer. An incorrect exclusion can eliminate a valid answer, so exclusion evidence must also be reviewed.

With filtering applied, six saved cases tested in three alternative orders produced 9 reference matches, 9 abstentions, and 0 incorrect concrete choices out of 18 runs. Because the alternative set changed, this cannot be converted into an accuracy improvement on the original questions. Good results on manually rewritten development questions were kept separate for the same reason.

In v9, I recorded adopted, overridden, or deferred after a judgment, and linked references to actual tool results and requirements. Adoption means the agent followed the answer. Tool success means the execution succeeded. Neither automatically proves that the answer was correct.

A primary model choosing an action after an abstention and completing the task does not make the abstention wrong. The review has to ask whether the packet contained enough facts at the time and whether more than one valid alternative remained.

If a batch contains one correct answer and one wrong answer, the whole batch must not be marked correct. Concrete choices and abstentions also need separate treatment. Follow-ups were linked to the original decision ID, rather than to another task's decision merely because it was the “latest call.”

That is also why v11 added bounded storage of reviewed inputs and answers on their owning runtimes. A hash can establish identity, but it cannot show what was missing or why an answer was wrong. Case retention was optional and available only when the caller had reviewed the material. Simple pattern checks were not presented as a guarantee that all secrets had been removed.

I also found installation tests mislabeled as ordinary work and corrected their provenance while preserving the originals. Including them in work-use or accuracy metrics would create the odd impression that more installation testing meant better task performance.

9. LAYA was fast, but I did not adopt it for this task set

To test another lightweight decision model, I temporarily installed LAYA on a 32GB Mac. The package was laya 0.3.21, and the checkpoint was convaiinnovations/laya-multilingual. It ran through MPS with Torch 2.8.0 and Transformers 5.17.0.

The installation files were about 659.9MiB. Model loading took 3.826 seconds, and the first inference took 2.133 seconds. Observed peak process RSS was about 2.64GiB, and MPS driver allocation was about 1.28GiB. Because the system used unified memory, I did not add those figures as though they were independent memory consumption.

For the first comparison, I took three actual tasks, constructed original questions and narrower factual questions, and used two alternative orders. That produced 12 requests per model, not 12 independent cases. Of the six narrower questions, LAYA matched the reference on 3 and KEV on 5.

Excluding the first call, median latency was 66.0ms for LAYA and 106.7ms for KEV. LAYA ran directly in the process, while KEV used remote HTTP, so this difference did not establish a difference in the models' computation speed alone.

In some cases, LAYA chose immediate replacement despite a prohibition on roster creation and deactivation, or declared an update complete before its first run had been confirmed. KEV also chose the wrong task scope after the alternative order changed, so neither side was free of problems.

For further testing, I wrote 23 development cases drawn from operations diagnosis, retries, roster preservation, access scope, and completion decisions, rather than shopping examples. Of those, 7 lacked enough evidence to establish a reference and were left unscored. One of the remaining 16 had lost an important condition during editing. I excluded it and treated the other 15 as the more conservative comparison.

15 development questions preserving important conditions, in three ordersLAYAKEV 0.8B
Reference matches2/4527/45
Concrete choices differing from the reference19/456/45
Abstentions24/4511/45
Execution failures0/451/45
Questions matching the reference in all three orders0/157/15

For latency across all 23 cases in three orders, LAYA's 69 successful requests had a median of 62.680ms and p95 of 118.192ms. KEV's 67 successful requests had a median of 134.6ms and p95 of 248.9ms, with two execution failures. The latency observations and the scored table above have different denominators.

These results do not establish LAYA's general performance. They test one multilingual checkpoint through its default execution path at the time, on a small task set I wrote myself. I did not calibrate thresholds separately or add training, and I did not test the other LAYA variants described in current documentation.

They were nevertheless enough to decide whether to adopt this configuration. It was fast, but its decisions did not preserve task conditions as well as I expected. I removed the temporary runtime and checkpoint, retained the result records, and did not add LAYA to production hooks.

10. Running PPLX 27B on a Mac

The next candidate was PPLX Decider v1 27B. As the 27B name suggests, its resource requirements were quite different from those of a small decision model. The original download target was 52,193,565,570 bytes, about 52.19GB. To reduce pressure on the shared connection, I limited the download to 5,000,000 bytes per second.

The target machine was an M4 Mac with 32GB of unified memory. Instead of loading the original weights unchanged, I ported the text path to MLX and converted the weights to affine 4-bit quantization with group size 64. The converted weight file was 14,420,466,057 bytes, about 14.42GB.

I did not connect it as a general conversation model generating long responses. I used its original decision readout to obtain a distribution over alternatives. The 255×5120 BF16 readout retained its values and dtype and was not quantized.

During the port, I compared a small four-layer FP32 model in Torch and MLX. At lengths 1, 17, and 256, the maximum absolute error was no more than 3.10e-6. Reloading the quantized weights and checking readout preservation also passed.

Numerical agreement on a small model, however, does not prove semantic equivalence between BF16 and 4-bit execution of the full 27B model. I did not run the full original BF16 model on the same cases, so I could not isolate the effect of quantization on decision quality. The original model's image path was neither ported nor tested in this experiment.

For the installation check, inference on a 368-token input took 6.267 seconds. Including model loading after module import, it took 12.417 seconds. Initially, each CLI process loaded the model afresh. Later experiments loaded it once and reused it.

Observed active model allocation was about 14.421GB. Further checks showed peak active allocation of 14.952GB, and the observed sum of active allocation and MLX cache after questions reached 21.474GB. These are not total operating-system memory figures, and they cannot simply be added to CPU RSS.

In some intervals, swap figures were unchanged before and after the experiment. Endpoint values do not prove there was no swap activity in between. The evidence supports short text requests running on 32GB, but does not guarantee stability with long context and multiple resident services.

11. Almost perfect at first—until I returned to original requests

PPLX and KEV received the same serialized packets. Their tokenizers and internal prompt formats differed, but the state, questions, and alternatives were kept the same. This comparison used the original candidate sets without exclusions.

First, I reused the 23 development cases written for the LAYA comparison. On the 16 questions with references, PPLX matched 16/16 in the original alternative order. Across original, reversed, and rotated orders, it matched 48/48. KEV matched 8/16 in the original order and 27/48 across the three orders.

Previously written development questionsPPLX 27B 4-bitKEV 0.8B
Reference matches in original order16/168/16
Reference matches across three orders48/4827/48
Abstentions across all question runs4/6918/69
Questions whose choices changed with order1/238/23

The results were good—and easy to overinterpret. Those 48 runs were the same 16 questions in three orders, not 48 independent problems. The development cases were human-edited questions, some already used in earlier tests. They were not a random sample of all work scored by an independent evaluator. The condition-loss case excluded from the conservative LAYA comparison was also included in this set of 16. The 48/48 result must therefore be read as reference agreement on those written packets.

I then checked whether the results held on original requests. I froze five previously submitted decision packets without rewriting their sentences or observations. One packet contained two questions, for six questions in total. Four had enough evidence to establish a reference.

Questions from original, unrewritten requestsPPLX 27B 4-bitKEV 0.8B
Reference matches in original order3/40/4
Reference matches across three orders9/124/12
Abstentions across all question runs3/182/18
Questions whose choices changed with order0/64/6

This original set also had selection bias: it deliberately included cases in which KEV had previously been wrong. It does not establish that KEV fails at this rate in ordinary work. PPLX's 3/4 is likewise not an estimate of general 75% accuracy.

PPLX distinguished choices about interference between concurrent execution and diagnosis, an existing retry schedule, and a directory-versus-file conflict. But when an authoritative roster differed from the web display, it abstained in all three orders. Preserving the roster was the reference under the original requirements, so the answers did not match. I kept that separate from choosing a concrete, incorrect modification.

Across the two sets, all 84 requests per model completed. That is an execution-success count, not a count of independent accuracy cases. I did not combine original requests, rewritten development questions, and unscorable questions into a single accuracy figure.

Adding harder variations

After the encouraging results, I created 12 stress variations from three original tasks, containing 15 questions. I included stale hypotheses, removed decisive observations, replaced current receipts with different contents, embedded instructions inside quotations, or extensively repeated irrelevant history.

Variations with altered receipts were hypothetical scenarios. They were not records of those events occurring in production, and the proposed actions were not executed.

Task-based stress questionsPPLX 27B 4-bitKEV 0.8B
Reference matches in original order12/158/15
Reference matches across three orders36/4524/45
Questions whose choices changed with order0/154/15
Abstentions12/4515/45
Matches on runs where abstention was the reference6/98/9

I reviewed the three items PPLX missed in the original order. It again abstained on roster preservation. When the decisive observation was removed, it selected a diagnostic action. When the condition was changed to a disabled retry, it abstained.

Two of these variations also had limitations in the questions themselves. The question with evidence removed still mentioned a particular diagnosis, while the alternative describing the absence of a retry also bundled in an additional scheduling action. I retained the mismatches against the pre-established references, but did not call all of them evidence of an actual incorrect action being executed.

Harder tests evaluate the quality of the problems as well as the models. Changing references after seeing model answers could improve the apparent results, so I kept the references and documented question limitations separately.

Quoted instructions and repeated context did not produce order-dependent choices in this small PPLX set. That is not evidence of general prompt-injection safety. Order stability, reference agreement, and security properties are separate measures.

12. Investigating why PPLX was slow

On the development questions, resident PPLX had a per-question median of 5.333 seconds and p95 of 7.246 seconds. KEV's full client median was 0.359 seconds, while its separately reported request-time median was 0.128 seconds.

Dividing 5.333 seconds by 0.359 seconds gives about 14.87 times. Against 0.128 seconds, it is about 41.8 times, but the timing boundaries differ more: one measures model question processing, while the other is the request interval KEV reports. My initial impression of a “difference of thousands of times” was not supported by these measurements.

Even so, adding 5~7 seconds to every decision is substantial. Keeping the model resident removed repeated loading but left the cost of computing the full input. A long input with repeated context, at 1,907 tokens, took about 31.5 seconds.

I first tested MLX compilation. With the same model resident, I alternated eager and compiled processing for two inputs, excluding the first execution of each form separately.

Input lengthEager medianCompiled median
278 tokens4.441 seconds4.469 seconds
432 tokens6.759 seconds6.755 seconds

I found no practical speed improvement. That means compilation did not help much in this implementation; it does not establish that this was the fastest possible implementation.

I also observed synchronized computation categories for the 432-token question. FFN took 4.860 seconds, linear attention 1.812 seconds, and full attention 0.523 seconds. FFN accounted for about 67.5% of the instrumented total of 7.198 seconds. An uninstrumented run took 6.894 seconds, so instrumentation boundaries affected scheduling and timing.

These were not GPU kernel traces or an analysis of maximum hardware performance. They still gave me a reason to investigate reducing the large model's input computation before removing a few Python function calls.

I also looked for a smaller PPLX sibling. The official offerings and related searches at the time did not reveal a smaller trained Decider checkpoint. A small general Qwen model would not automatically become the same decision model. Both an appropriate base-model size and training for the decision head would be needed.

Being published on HF did not prevent customization. I checked the public training configuration and model implementation. I had already ported and quantized it, but did not attempt training or distillation into a smaller model. The public recipe also did not include the exact training data in full, so I cannot claim the same model could be reproduced from it.

13. Reusing shared state across different questions

My first question about caching was how often the same question and answer would actually recur. Operations agents ask different questions each time. A cache storing only identical answers might offer limited value.

Here, I tested reuse of shared-state computation rather than an answer cache. If separate questions ask about retries and file conflicts from the same observation set, their questions differ but their leading state is the same. Shared-prefix reuse already exists in inference systems. The focus here was adapting it to this decision model's local text path.

I checked for an exactly matching state prefix and stored its computed state. For each question, I copied the stored state into an independent branch. This model path has both recurrent state and a KV cache, so I avoided passing the state modified by one question directly to the next, or merely rolling back the token length.

I also did not assume that apparently identical JSON state guaranteed an identical token prefix. I checked full-input tokens and prefix boundaries because tokenization can be affected by the boundary. When the model or state changes, the old entry is not reused.

공통 상태 prefix 재사용 같은 상태의 정확한 토큰 prefix를 계산해 보관하고 각 질문에 독립 복사한다. 질문과 답이 같을 필요는 없지만 상태 또는 모델이 바뀌면 다시 계산한다. 질문이 달라도 관찰한 상태는 같을 수 있다 같은 관찰 · 요구사항정확한 prefix 경계 확인모델·상태 변경 시 무효화 상태 계산 결과를 한 번 보관recurrent state + KV cache완료된 답을 그대로 재사용하는 방식과 구분 질문마다 보관 상태를 독립 복사 질문 1: 재시도 상태분기 1 계산 → 답 1 질문 2: 파일 충돌분기 2 계산 → 답 2 새 상태의 첫 계산 비용은 남는다 · 확률 동등성의 미통과 사례도 남긴다
Independent question branches from the same state prefix

The following request times were measured with the model resident. They include tokenization, cache creation on a miss, branch copying, and all questions in the request.

Input drawn from workQuestionsUncached requestCache missCache-hit median
Interference between concurrent diagnostic runs16.763 seconds6.814 seconds2.427 seconds
Retry state and file conflict213.123 seconds8.878 seconds4.058 seconds
Authoritative roster versus display mismatch14.676 seconds5.206 seconds2.079 seconds
Artificially repeated historical records131.467 seconds31.975 seconds2.649 seconds

Follow-up questions on the same state reduced waiting time. A request containing two questions shared computation even on the initial miss, reducing approximately 13.1 seconds to 8.9 seconds.

For a single question with fresh state, however, cache creation was not free. A miss was similar to or slower than uncached execution. Saying “PPLX now takes 2 seconds” would be inaccurate without the state-reuse condition.

These tests ran uncached execution first and cached execution afterward. They were not production benchmarks with randomized execution order and system load. Cache-hit medians for task inputs were related observations from three alternative orders, and the long repeated input was an artificial stress case.

The choices matched, but the probability distributions did not match exactly

All selected answers matched across 36 cached-question comparisons. Two comparisons, however, exceeded the pre-established tolerance of 0.005 for differences in candidate probabilities. The maximum difference was 0.009574, about 0.9574 percentage points.

Small FP32 cache checks and branch-independence checks passed, but probability equivalence for the full quantized model did not pass every comparison. Splitting computation may have changed low-precision numerical results, but I did not establish that as the cause.

This result would be easy to hide because every final choice matched. Near a decision boundary, however, a changed distribution could change the choice on another input. I therefore did not describe this implementation as completely lossless.

Stored prefix arrays were about 171~271MB. That excludes model weights, per-question copies, allocator overhead, and operating-system costs. The current implementation stores only one prefix. Alternating tasks could reduce hits, but I have not measured the actual production hit rate.

Shorter inputs were a separate test

Manually rewriting state while preserving important observations and relationships improved uncached latency on ordinary task inputs by about 0.2~6.9%. On the artificial repeated input, reducing 1,907 tokens to 409 tokens changed 31.467 seconds to 6.752 seconds.

The large difference mainly came from removing deliberately repeated material. It does not support a claim of more than a 4× speedup in ordinary work. I also did not measure the time spent rewriting state or the cost of preserving meaning in actual work.

Shorter evidence is not always better. Remove one important condition and the model can make a wrong decision faster. The development case excluded from the LAYA comparison because it lost a condition was a reason to check information preservation alongside speed.

14. Removing duplicate tokenization was a small improvement

Reviewing the resident service's input flow showed that it tokenized full questions to check submission limits, then tokenized the same questions again inside the inference function. The earlier stage did not pass its validated token lists to the next stage.

I repeatedly measured eight actual development inputs on the CPU. Median additional tokenization time for ordinary inputs was 0.255~0.668ms; the long repeated input took 1.421ms. Initial tokenizer loading was a separate cost.

That is tiny compared with seconds of model computation, but there was no reason to do the same work twice. I changed the service to pass per-question-ID token lists from admission validation into inference. The original tokenization path remained available when calling the inference function independently.

The service's limit on the sum of tokens across all questions remained. Tests covered accepting 4,096 tokens and rejecting 4,097, rejecting mismatched question IDs, and not running inference after preparation failed. I did not claim to remove tokenization needed to verify prefix boundaries.

In two before-and-after HTTP tests, choices, token lengths, prefix lengths, and probabilities matched, with a maximum probability difference of 0.0. I also compared hashes of the deployed and tested sources. An actual over-limit input returned a limit error without invoking the model.

This removed duplication in our service wrapper; it was not a claim to have found a bug in the Hugging Face original. The generalizable pattern is to retain tokens already produced during validation through to inference. I did not make a controlled remeasurement of the reduction in full request time, so I assigned no separate large speedup figure to this change.

15. v12 receives two answers in parallel

To observe both 0.8B's speed and 27B's decision quality on this task set, I changed v12 to call both models concurrently. It validates the decision input once, applies explicit exclusions, and sends the same packet to KEV and PPLX.

The existing answers field remains KEV's answer. PPLX's answer goes into comparison.answers, with disagreements recorded per question. If they differ, the primary model checks the evidence and requirements. Agreement does not establish correctness either.

v12의 같은 판단 패킷 병렬 호출 보이는 근거로 만든 패킷을 검증하고 같은 패킷을 KEV와 PPLX에 병렬 제출한다. 답은 따로 보존하고 근거 검수 후 허가된 행동과 후속 기록으로 연결한다. PPLX가 바쁘거나 사용 불가여도 KEV를 보존한다. 요청 · 현재 관찰 · 요구사항 · 미확인 사항 하나의 명시적 판단 패킷로컬 검증 → 출처 있는 제외 적용 → 동일 패킷 제출 KEV 0.8Banswers기존 기본 답을 보존 PPLX 27B 4-bitcomparison.answers바쁨·실패는 별도 진단, KEV 유지 주 모델이 두 답을 각각 근거와 대조동의 ≠ 정답 · 다수결 없음 · 자동 우선권 없음 허가된 행동 → 도구 결과 → 원래 판단의 후속 기록
v12 parallel calls with the same validated packet and separate review

This parallel setup is not a race that returns whichever answer arrives first. On the normal path, it waits for the slower model so the answers can be compared. Overlapping calls can reduce time relative to sequential execution, but the setup does not preserve KEV-only latency.

In simplified terms, total latency includes packet preparation and validation, the longer of the two model processing times, and answer review and action. The implementation also involves communication and scheduling. Adding two decision models does not eliminate an equivalent amount of the primary model's reasoning time.

For a short HTTP-status question after installation, the first state took 2.829 seconds overall: KEV took 0.435 seconds and PPLX 2.808 seconds. A request reusing the same state took 1.141 seconds overall: KEV took 0.176 seconds and PPLX 1.122 seconds.

These were installation checks, not medians across ordinary task questions. Earlier first questions on state drawn from actual work took 5~7 seconds or more. Follow-up questions using stored state took roughly 2~2.4 seconds, and multiple questions could take longer.

What happens when it is busy?

The PPLX service has one inference worker and one stored prefix. If a new request overlaps an inference already in progress, it rejects the request with comparison_busy. KEV's answer is retained in that situation. In a controlled overlap test, KEV's answer arrived in 0.155 seconds, alongside a PPLX busy diagnostic.

I did not add automatic retries. PPLX has a 75-second response timeout, and model computation may continue to occupy the worker even after the HTTP wait ends. The limit on the sum of question tokens is a conservative admission limit to reduce that burden. It is not the same as the model's theoretical maximum context length.

I applied the v12 client and instructions to five Codex runtimes and six Hermes runtimes and performed installation checks. Those are runtime-profile counts, not eleven computers. Whether every existing GUI session reloaded its settings, and whether older bundled fallback clients use the same path, remain separate checks.

Lifecycle hooks are still checkpoints that do not call models. Both models are called when an actual, concrete decision packet is submitted. Losing that distinction could create an expensive setup that calls 27B on every tool use, or reports thousands of checkpoints as actual decision consultations.

16. What actual monitoring showed

In the last daily observation window before v12, there were 79 actual consultation records, 98 questions, and 38 abstentions. The question-based abstention rate was 38.8%. I separately excluded 240 telemetry rows with installation or simulation provenance. That excluded row count does not mean 240 consultations.

This window recorded v11 client use. It cannot be presented as a measurement of v12's parallel-comparison effects. Calls for experiment management and monitoring can also be mixed into actual consultation records.

In the preceding window, 11 of 60 questions were abstentions, or 18.3%. It might be tempting to connect the early installation-heavy 92.7% and later figures in the 30~40% range into a graph showing steadily improving accuracy. But question content, denominators, window lengths, versions, and task mixes changed. Fewer abstentions could also mean more incorrect concrete choices.

The final window had no recorded decision failures or follow-up recording failures. Follow-ups were linked to 73 of 79 consultations. Action references existed for 64, and requirement references for 54. Additional tool-call counts and rework counts were each recorded for only seven consultations.

Item in the final v11 observation windowValueWhat it supports
Consultations / questions79 / 98Actual consultations and questions in the logs
Question abstentions38/98, 38.8%Fraction of answers that were abstentions
Linked follow-ups73/79Fraction with follow-up records
Action references64/79Fraction with references pointing to execution results
Requirement references54/79Fraction with requirement references
Tool-call and rework counts7/79 eachVery limited observation coverage for cost comparisons

Across runtimes with active consultations, median request latency ranged from 140.6~258.0ms and p95 from 301.9~409.8ms. These runtime metrics mixed different tasks, so I did not turn that range into a single combined p95. Short checkpoint execution times were excluded from consultation latency.

I also selected up to five visible completed tasks and checked their normal receipts and results. I looked at concrete questions: whether a transfer involving only two targets preserved that scope, whether an existing workflow was reused, and whether usage observations alone had led to a skill change.

There was evidence connecting some choices to actual outcomes. Checking normal receipts, however, is different from independently auditing all remote data again. I did not measure whether another choice would have been worse. Successful tasks do not turn every abstention into a wrong answer.

Most importantly, there were no comparable baselines for full task time or tokens. Fast calls were verified, but reductions in total work time or token cost were not yet proven. The single early comparison with the existing skill cannot serve as a baseline for every different later task.

17. How the criteria for delegating decisions changed

At first, it was easy to assume that a fast decision response would make the primary model faster. In practice, constructing the question, interpreting the answer, and handling mistakes accounted for substantial work.

The current system keeps the following rules.

Start from observed facts. Include the request's exact conditions, current tool results, and known unknowns. Do not choose a conclusion first and ask for agreement. Distinguish past and current state, and a running process from verified functionality.

Give each question one decision. The existence of a retry and the need to fix a file conflict can be separate questions. Batch independent questions sharing evidence, but do not treat a later question that depends on an earlier answer as independent in advance.

Describe actual actions in the alternatives. Do not hide meaning in IDs or favorable adjectives. Use source-backed exclusions for explicitly forbidden actions. Preserve uncertainty.

Check answers against the evidence again. Real cases changed their answers when alternative order changed. In ordinary work, I do not repeatedly reorder alternatives and take a majority vote. A consultation after an abstention is allowed only on a limited basis when there is genuinely new evidence or corrected question scope. Success on a narrower new question is not scored as improvement on the original.

Separate authority from judgment. A decision-model choice is not new authorization for deletion, payments, security changes, or permission changes. Existing user instructions and runtime limits still apply. A fast auxiliary model does not expand authority.

Return to the original decision after acting. Link adoption, rejection, or deferral to actual tool results, and record unverified matters as unknown. Evaluate successful outcomes separately from correct judgments. When comparing models, retain separate verification labels for KEV and PPLX.

Written this way, the rules may seem obvious. In the actual integration, places that omitted them produced distorted metrics or failures with unclear causes.

18. What is established, and what remains unverified

The small models were clearly fast. KEV 0.8B at the time answered many concrete questions within hundreds of ms. LAYA was also very fast on the direct execution path tested. Speed alone did not mean that task conditions were respected.

PPLX 27B 4-bit had better reference agreement and alternative-order stability than KEV in this task-based comparison. But the easy development set's 48/48 was followed by abstentions on original requests and mismatches on stress cases. “Effectively 100%” can apply only to that particular development set.

The shared-state cache helped when its conditions held. It worked across different questions on the same state, and batching questions shared computation on the first request. Fresh state, cache competition, long inputs, and small probability differences remain issues.

Removing duplicate tokenization was a small, properly checked repair. It was useful technical cleanup, but not something to present as a change in seconds-long inference. I also retained the compilation results as observed. An optimization that did not make execution faster still helped narrow the next choice.

The next evaluation of benefits needs a comparable task set more than another new version. There are three areas to check.

First, original packets and requirements should be frozen in advance, with an independent evaluator setting references and unscorable conditions where possible. Development cases, selected failure cases, and ordinary actual work need separate counts. Evaluation should cover incorrect concrete choices and useful decision coverage alongside abstention rates.

Second, each full task should be measured once, from its start through outcome verification. That includes preparation, decision calls, primary-model review, additional tool calls, and rework, compared with a baseline task of the same kind. Summing response times or multiplying cache-hit speed does not establish total savings.

Third, actual state reuse and PPLX contention under overlapping work need measurement. With the current single prefix and single worker, that means hit rates, rejection rates, p50/p95 for fresh and cached state, and outcomes when proceeding with KEV alone. These measurements cannot be described as already complete.

The chosen v12 setup is an experiment that observes a fast small model and a heavier comparison model on the same input, with the primary model checking the evidence. It is not an automatic arbiter of correctness. Whether it makes whole tasks faster and cheaper remains to be tested.

Still, the original question now has a more concrete answer. Fast decision models can be attached to agents. To make them useful, the decision problem must first be constructed properly, and the answer connected to execution and outcomes. The failure records and measurements under specific conditions were more useful for choosing the next implementation than any single best-looking number.

Appendix A. Configurations needed to interpret the tests

These conditions matter when interpreting the experiments in this article. The table is not a promise that installing a new version under the same product name will reproduce the results.

CandidateConfiguration used at the timeConditions to keep in mind
KEV 0.8Bjaredpalmer/kev-0.8b, pinned revision, existing local serviceFresh versus repeated inputs; different HTTP and client timing boundaries
KEV 4BThe then-current kev-4b, MLX BF16, including a separate 32GB Mac testKeep timeouts on the smaller Mac separate from successful runs
KEV 8BHistorical kev-8b, Torch 2.8.0 / Transformers 5.17.0, MPS BF16Distinct from other sizes in the current list; different backend from 0.8B
JEVPinned jev-1.13.0, hosted APIClient-plus-communication times for six requests; server internals and billing unverified
LAYAlaya 0.3.21, laya-multilingual, direct MPS executionWritten development cases; other variants, training, and calibration untested
PPLXpplx-decider-v1-27b, MLX affine 4-bit group 64, BF16 readout preservedM4 32GB; text only; no full BF16 or image-path comparison

The PPLX port used MLX 0.32.2 and MLX-LM 0.31.3. API details were checked against installed source and documentation. Model files and inputs were checked for fixed identity, but these experiments were not a unified model benchmark on the same hardware, backend, and timing boundaries.

Key checkpoint identifiers were:

KEV 0.8B: 9a45d25eb2ab761841196625383fa1dff0e56c1e
Historical KEV 8B: c80773da7f383f93c4dbff0c0b008e0463f9145a
LAYA multilingual: e4e9ddf21a7b1903b7acffd8814ad4307bf63a67
PPLX Decider: 5117a6c7fe73b19308dc1a6b0fb529a40c2ecad4

The configuration table and fixed-input methods explain experimental conditions, but the article itself is not a public reproduction package.

Appendix B. Information the decision packets were intended to preserve

The actual client performs more validation, but this is the core shape of a choice packet. The example illustrates the structure; it is not a scored experimental case.

{
  "state": {
    "observations": [
      "The retry job is active and its next run is scheduled.",
      "The export target must be a directory, but a regular file occupies that path."
    ],
    "requirements": ["Preserve the existing scheduled job."],
    "unknowns": ["Actual export success after resolving the conflict has not yet been verified."]
  },
  "questions": {
    "first_inspection": {
      "type": "choice",
      "instructions": "What should be inspected first to resolve the export failure?",
      "criteria": {
        "path_collision": "Inspect the file-versus-directory conflict at the export target.",
        "retry_absence": "Check whether the retry schedule is missing.",
        "insufficient_evidence": "The supplied observations do not distinguish the two inspections."
      }
    }
  }
}

Schema validation checks this JSON's shape and allowed fields. It does not automatically establish factual sufficiency, neutral alternatives, or authority to execute an action. Linking the packet to original evidence and reviewing the answer remain necessary.

Appendix C. Units distinguished when reading the numbers

References

The figures in the article's tables come from the local experiments and operational observations described here. They were not mixed with the projects' official benchmark figures.