Published on

Deciphering and Scaling Jev

Authors

Introduction

Jev is a model created by TypeSafe AI, that is currently taking the developer community by storm. It makes lightning fast structured decisions while costing only a fraction of what frontier LLMs would charge. TypeSafe AI has made it clear that Jev is not an LLM but rather part of its own class of models called System One Models. With little else stated regarding the internals, this launch has expectedly raised several questions over architecture, capabilities, speed, etc. Using my expert sleuthing skills (and a lot of claude) I've decided to decipher what I can about how the model works and I am sharing my findings in this article.

Take everything I say with a grain of salt as it is all speculation and let's dive right in.

thumbnail

What do we know about Jev

Unlike LLMs Jev has no auto regressive capabilities, instead outputs are determined by a fixed schema defined in the input. The input consists of 3 main components: state, question(s), and criteria.

State

The state is the context of the environment, it can be a ticket, a json object, an article snippet, or anything that can be represented by text. There is only one state per request.

Questions

questions describe what we want to know about the state. Jev currently supports three types.

  1. Choice question: the user defines choices and the model returns a probability distribution over all of them. This is representative of how close the model thinks a choice is to the provided state.

  2. Score question: levels are defined in the schema and the model is in charge of determining which level the state best matches. The score determined by the model may not be a whole number, thus a score may straddle 2 levels.

  3. Noul question: A binary classification problem, where the result is between 0 and 1. 0 means No and 1 means yes.

Criteria

This defines what the question is evaluating. For Nouls this is what yes and no mean, for choice this is what each choice represents, and for score what each level entails.

Beyond this there are practically no other requirements. You can ask as many questions as you want per request, in any combination, with the only ceiling being the model's context window. Here is a sample request taken straight from the docs:

Sample request
{
  "state": {
    "source_text": "Invoice #4471 issued March 3, 2026 to Beaver Dam Logistics for $12,840.00, net 30."
  },
  "questions": {
    "invoice_number_is_correct": {
      "type": "noul",
      "instructions": {
        "field": {
          "name": "invoice_number",
          "type": "string",
          "description": "The identifier printed on the invoice."
        },
        "extracted_value": "4471",
        "question": "Does `extracted_value` match the `field` as it appears in `source_text`?"
      }
    },
    "customer_name": {
      "type": "choice",
      "instructions": {
        "field": {
          "name": "customer_name",
          "type": "string",
          "description": "The organization the invoice was issued to."
        },
        "question": "Which option is the value of `field` in `source_text`?"
      },
      "criteria": {
        "Beaver Logistics": null,
        "Dam Logistics": null,
        "Beaver Dam Logistics": null,
        "Beaver": null,
        "Dam": null
      }
    },
    "amount_due": {
      "type": "score",
      "instructions": {
        "field": {
          "name": "amount_due",
          "type": "number",
          "unit": "USD",
          "description": "The total the invoice asks to be paid."
        },
        "question": "How large is the `field` value in `source_text`?"
      },
      "criteria": [
        "Under $1,000",
        "$1,000 to $10,000",
        "$10,000 to $100,000",
        "$100,000 to $1,000,000",
        "Over $1,000,000"
      ]
    },
    "payment_terms": {
      "type": "score",
      "instructions": {
        "field": {
          "name": "payment_terms",
          "type": "integer",
          "unit": "days",
          "description": "Days allowed for payment, from terms such as \"net 30\"."
        },
        "question": "How many days does the `field` in `source_text` allow for payment?"
      },
      "criteria": [
        "Due on receipt",
        "Net 10",
        "Net 30",
        "Net 60",
        "Net 90"
      ]
    }
  }
}

The big thing we should take away from all this is that Jev is essentially a robust classification model, and from that we can derive some solid ideas on the architecture.

Model Architecture

Now the only architectural hint TypeSafe tells us is that Jev is not an LLM. It's important to understand that not being an LLM and not using a transformer are 2 different things. In fact there are plenty of models that rely on a transformer block that don't generate text (ViT, AST, etc.). In my opinion stating that Jev is not an LLM purely means that Jev is not a decoder based model, leaving room for the possibility that Jev is probably some spin on an encoder.

In particular I think it's an encoder that uses a hierarchical mask for attention, along with a unique scoring head (as opposed to an LM head). To prove my point let's start with the attention mechanism.

Bidirectional attention with hierarchical masking

Decoder based models use a causal mask (a triangular mask) when doing attention, this is a consequence of the fact that LLMs are token generators, and that present tokens cannot pay attention to tokens that do not exist yet.

Causal mask visualization

Jev on the other hand does not generate tokens, it returns scalar values that get reinterpreted to actions. As there will never be future tokens, it makes no sense to use causal attention, which leads me to the conclusion that like an encoder it uses bidirectional attention.

Bidirectional mask visualization

Now this doesn't mean that masking isn't present. The corresponding documentation and blogs consistently mention that Jev is designed to process multiple queries together

Parallel. Generates all outputs in a single query.

While also hinting in several places that each classification is evaluated independently (without respect to other levels, classes, ...).

Every level is evaluated separately. The model doesn’t see a level’s number or its neighbors

Now for all questions in a request to be processed in parallel, batching of the questions must occur (most likely concatenated across the sequence dimension). And to ensure independent classification, attention leaking from one classification to another must be prevented.

This leads me to believe a hierarchical masking pattern is used where each component in the tree can pay attention to itself and its parents but not its children or siblings.

Hierarchy Tree Hierarchy qk

Scoring Head

In an LLM the LM head layer is in charge of generating the logits. We get them by multiplying the hidden state by the embedding weights returning a Matrix of size [sequence_length, vocab_size].

Hierarchy qk

Because Jev does not generate tokens, using the embedding weights as the weights matrix makes no sense. Instead we want a single scalar value which through transformations we can then use as our output. To achieve this Jev probably does a simple dot product between the hidden state ([num_of_classifiers, embedding_size]) and a score vector ([embedding_dimension]) or they employ a small MLP.

Hierarchy qk

How is it faster than an LLM?

Compute Bound

If you are familiar with LLM inference, you are aware that it is separated into 2 distinct stages: Prefill and Decode. Prefill is everything involved with the generation of the first token, decode is everything after, up to the end of the turn. Prefill is compute bound assuming a long enough context is passed in, while decode is always memory bound.

As Jev does not generate tokens it never has to enter the decode stage. This is the most time consuming part as all the active weights have to be streamed from slow HBM for every token generated.

This also explains the claim made here:

Incredibly efficient and hardware-aware.

Since Jev's forward pass is basically only prefill, it better utilizes the GPU(s) (computation heavy work).

Now it's impossible to claim that this is the sole reason for the speedup as it's heavily determined by the number of tokens generated by the competing model, but it's certainly one of the main ones.

Model Size

The only other major explanation for the speed is model size, while the blog directly claims the model is not small, I doubt the model reaches frontier scale. We can estimate the model size using the pricing that was provided.

First let's establish some assumptions:

  1. Cost: 0.042 per million tokens→4.2×10−8 per token0.042 \text{ per million tokens} \rightarrow 4.2 \times 10^{-8} \text{ per token}
  2. FLOPs per input token: ∼2N\sim 2N (N being number of parameters in the model).
  3. Cost per FLOP:
DType~ $ / hourFLOP / sCost Per Flop
H200FP82.202 PFLOPS~3.1 × 10⁻¹⁹ $
B300FP84.505 PFLOPS~2.5 × 10⁻¹⁹ $
B300FP44.5015 PFLOPS~8.3 × 10⁻²⁰ $

When solving for the parameter size with break even pricing we get:

2N⋅Cost Per FLOP=4.2×10−8  ⟹  N=4.2×10−82⋅Cost per FLOP2N \cdot \text{Cost Per FLOP} = 4.2 \times 10^{-8} \implies N = \frac{4.2 \times 10^{-8}}{2 \cdot \text{Cost per FLOP}}

DTypeN
H200FP8~69B
B300FP8~84B
B300FP4~252B

Now this is the most unrealistic case: it assumes that the model is achieving peak utilization and that TypeSafe is ok with just breaking even. Adjusting our results to assume 50% efficiency at a 50% profit margin puts N at a quarter of what was displayed.

Which is why I estimate the model is roughly 17B ~ 63B parameters.

Inference Scaling

With all these assumptions in mind let's share how I'd scale Jev for inference and how it differs from traditional LLM inference.

N/A

The big LLM inference levers of Speculative Decoding, Caching (KV and Prefix), Paged Attention, and disaggregation, don't really apply here.

  1. Speculative decoding is contingent on token generation and a decode phase.
  2. KV caching only makes sense in a regime where a decode phase is present.
  3. Prefix caching seems like the one you would reach for, since questions should remain static. However 2 things block it:
    • First, the questions attend to the state and the state likely changes per request.
    • Second, the attention is bidirectional meaning if one token changes the entire projection changes.
  4. Paged attention assumes you have a kv cache which we should not. Regular flash attention is something we should use but adapted to handle the hierarchical mask.
  5. Disaggregation (prefill decode split) needs a decode phase.

Scheduling / Batching

Unlike LLMs which can leverage large batch sizes during decode, Jev can only batch until the compute is saturated. This is because once a batch is large enough to saturate the compute, additional batching will simply scale the total time required to process the batch.

  • For short context this has no real effect. We can leverage this to pack requests together up to the saturation point.

  • For long context this poses a serious problem as now requests are essentially serialized.

Model Sharding

DP

The easiest way to fix this scaling problem is to simply have more models, with data parallel we increase request throughput with each new replica we add.

DP

TP

If request latency becomes a problem we can utilize tensor parallelism to shard layers (attention, dense matmul) across multiple GPUs, lowering the total compute done by a singular GPU and forward pass time. This however needs to be balanced with the added communication overhead.

TP

In production a combination of both may be used, here is an example of DP 2 with TP 4:

All P

Conclusion

I don't think the magic of Jev lies in the architecture, while it's true the model is extraordinarily fast, it becomes obvious why that's the case with some prior inference knowledge. Speed is a byproduct of the purpose of the model, while the real achievement lies in how it was trained and what it is able to do. I've worked in the ML space when classification was all the rage, and seeing a model which generalizes that very problem is extremely cool. Hopefully this is only the start of System One Models and I'm very excited for what the future holds.

Kudos to the TypeSafe team, and feel free to reach out if you have any suggestions/criticisms/thoughts on the article.