Four answer types.
Yes/no, pick one, rate a result, or select all that apply. Ready to use in your application.
Turn context into decisions.
A small AI model that reads once, answers many,
and runs on infrastructure you control.
“I was charged twice for my subscription. Please refund the duplicate payment.”
Needs a refund?
Which team?
Which topics?
Accuracy on our project’s test questions, not a guarantee for every use case. What did we test? ↗
Route a request. Check an agent’s work. Apply a policy. SelfJev turns your text and questions into structured decisions with probabilities—without waiting for a generated response.
Yes/no, pick one, rate a result, or select all that apply. Ready to use in your application.
Run the model on your own GPU. Process sensitive text without sending it to an external model API.
Already using Jev? Keep the same request format and point your application to your own server.
A support message can raise several questions: what happened, who should handle it, and what to do next. SelfJev reads the message once and reuses that work for every answer.
SELFJEV / SHARED-PREFIX TREE
01. Give it contextA message, document, or AI response, together with the questions you need answered.
02. Ask several questionsEach question uses the same text, so the model avoids reading it from scratch each time.
03. Act on the answersGet clear choices and probabilities to route, filter, or review in your own application.
Your documents, your vocabulary, your edge cases. Control the context budget and adapt the model to the decisions that matter to you.
Jev allows 32K tokens for your text plus the longest question. SelfJev has a configurable limit, built on a model with a 262,144-token native context window.
SelfJev defaults to 32,768 tokens; larger windows need sufficient memory and validation. The full base-model window has not been validated in our engine. Jev also allows 64K across a whole request.
Context & sizingNot getting the decisions you need? Fine-tune SelfJev on examples from your own workflow: your labels, your policies, your definition of a good answer.
Can it choose the right answer from a piece of text? We tested yes/no decisions, choosing between options, and selecting every answer that applies.
How often did the model choose the expected answer? The same 1,991 questions test yes/no decisions, choosing one answer, and selecting all that apply.
Scores are the share of answers matching the expected result; for “select all,” every choice must match. Only models with a recorded result are shown. Click a row for its report.
SelfJev approaches Jev’s accuracy on these tests and runs on your own infrastructure. The questions and expected answers were created and checked by AI; results on your own data may differ.
How we tested itHandling exceptions, weighing alternatives, and checking AI responses all needed targeted training. Improving those examples mattered more than simply making the model bigger.
From a local Mac to a dedicated GPU. Explore measured response times, adjust your workload, and see what fits your infrastructure.
H100 · measured processing time
Very short input · one question
The complete merged weights, with tokenizer and model configuration.
Get the model ↗LORA ADAPTERThe trained adapter used by the native tree engine, with its model card.
Get the adapter ↗EVALUATION DATASET3,657 questions across text decisions, AI response review, and record reasoning.
Explore the dataset ↗A lightweight Python SDK. One GPU to serve. Docker and an AWS deployment command, with practical guides for Runpod and Google Cloud.
from selfjev import SelfJev, Noul, Choice
# Use the secret you set as SELFJEV_API_KEYS on your server.
client = SelfJev(
base_url="http://localhost:8000",
api_key="your-server-key",
)
result = client.system_one(
state="I was charged twice. Please refund me.",
questions={
"refund": Noul("Does the customer want a refund?"),
"team": Choice("Which team should handle this?", {
"billing": "payments and refunds",
"support": "technical support",
}),
},
)