Typesafe system one models and new /v1/decisions endpoint
Bifrost now supports TypeSafe's jev judgment models behind a unified POST /v1/decisions endpoint and a 1:1 /typesafe drop-in for TypeSafe's official SDKs. This is available in v2.2.2 for both OSS and Enterprise.
TypeSafe's jev family takes a piece of state - a support ticket, a model's answer, a review, any string or JSON blob - and returns calibrated judgments: a probability, a pick from your options, or a rubric score, each with confidence and a full probability distribution. That makes them a natural fit for the decision points scattered through every AI product: content moderation, ticket routing, eval-in-the-loop grading of LLM outputs, churn and escalation triage, guardrail checks. Anywhere you force a chat model to answer "just say yes or no", a judgment model does it faster, cheaper, and with a number you can actually threshold on.
TypeSafe support, two ways in
Bifrost now supports TypeSafe's System One models (jev-1.13.0, jev-latest, jev-preview) as a first-class provider, with a new operation type built for judgment workloads.
POST /v1/decisions is the unified Bifrost format. Send state plus a map of named questions - noul for probabilities, choice for classifications, score for rubrics - and get one typed answer per question, with confidence, probabilities, and legend metadata normalized into a single value field. Model routing, virtual keys, governance, logging, and cost tracking all apply, and every request is strictly validated before it costs you anything.
The /typesafe integration is the 1:1 drop-in. If you already use TypeSafe's Python or JavaScript SDK, point TYPESAFE_BASE_URL at your Bifrost deployment's /typesafe prefix and you are done - one environment variable, zero code changes. Native request and response shapes are preserved byte-for-byte on the success path, and Bifrost adds what upstream does not have, including a models endpoint and automatic retries on rate limits and overloads.
LLM fallbacks
If jev is down or for some reason API calls fail - we have added support for LLM fallbacks on both /v1/decisions and /typesafe/v1/systemone endpoint.
How it maps to a normal fallback call
A decision request already carries a fallbacks list, exactly like chat and responses requests. The trick is what happens when a fallback provider has no native decision endpoint (only typesafe does today). Instead of returning "unsupported", Bifrost emulates the decision through that provider's Responses API:
- The decision dispatch calls
provider.Decision(...). - If the provider returns unsupported_operation (every non-typesafe provider), one shim at the dispatch layer takes over - no per-provider code.
- The shim rebuilds the question set as a single forced
emit_decisionfunction tool (tool_choice: "required"), sends it via provider.Responses, and maps the tool-call arguments back to the exact DecisionResponse shape - values, per-answer confidence, choice probabilities, and score legends.
Because emulation runs through the same fallback loop, both the primary path (name an LLM as the decision model) and the fallback path (LLM after jev fails) flow through that one shim.

A decision request already carries a fallbacks list, exactly like chat and responses requests. The trick is what happens when a fallback provider has no native decision endpoint (only typesafe does today). Instead of returning "unsupported", Bifrost emulates the decision through that provider's Responses API:
- The decision dispatch calls
provider.Decision(...). - If the provider returns
unsupported_operation(every non-typesafe provider), one shim at the dispatch layer takes over - no per-provider code. - The shim rebuilds the question set as a single forced
emit_decisionfunction tool (tool_choice: "required"), sends it viaprovider.Responses, and maps the tool-call arguments back to the exact DecisionResponse shape - values, per-answer confidence, choice probabilities, and score legends.
Because emulation runs through the same fallback loop, both the primary path (name an LLM as the decision model) and the fallback path (LLM after jev fails) flow through that one shim.
We did a 30 question benchmarking keeping jev as ground truth to ensure it doesn't impact the quality of the output. We have picked a few models in random order as an example : overall testsuite runs across 100+ models.
{
"model": "typesafe/jev-1.13.0",
"fallbacks": ["openai/gpt-5.6-luna", "anthropic/claude-sonnet-5"],
"state": "...",
"questions": { "urgency": { "kind": "score", "instructions": "...", "criteria": ["low","medium","high"] } }
}If jev is reachable, jev answers. If not, gpt-5.6-luna emulates the same decision and returns the identical response shape - the caller sees no difference beyond the resolved model name.
Client request (jev primary, gpt-5.6-luna as fallback)
{
"model": "typesafe/jev-1.13.0",
"fallbacks": ["openai/gpt-5.6-luna"],
"state": "Customer: you charged me three times for one order, refund today or I cancel.",
"questions": {
"urgency": { "kind": "score", "instructions": "How urgent?", "criteria": ["low", "medium", "high"] }
}
}jev is down, so the request falls to openai/gpt-5.6-luna
Bifrost catches the failure, moves to the fallback, and since openai has no native decision endpoint it emulates the decision through the provider's Responses API. This is the call it actually makes to gpt-5.6-luna:
POST (internal) provider.Responses -> model: gpt-5.6-luna
{
"instructions": "You are a judgment engine. Read the state and answer every question by calling the provided function exactly once...",
"input": [{ "role": "user", "content": "Customer: you charged me three times for one order, refund today or I cancel." }],
"tool_choice": "required",
"tools": [{
"type": "function",
"name": "emit_decision",
"parameters": {
"type": "object",
"required": ["urgency"],
"properties": {
"urgency": {
"type": "object",
"required": ["value", "confidence"],
"properties": {
"value": { "type": "number", "description": "... Score levels (index -> meaning): 0=low, 1=medium, 2=high" },
"confidence": { "type": "number", "minimum": 0, "maximum": 1 },
"probabilities": { "type": "object", "additionalProperties": { "type": "number" } }
}
}
}
}
}]
}gpt-5.6-luna is forced to call emit_decision, and Bifrost maps its tool-call arguments back to the standard decision shape.
Response the caller receives (identical shape to jev)
{
"model": "gpt-5.6-luna",
"answers": {
"urgency": { "kind": "score", "value": 2, "confidence": 0.95, "probabilities": { "0": 0.0, "1": 0.02, "2": 0.98 }, "legend": { "0": "low", "1": "medium", "2": "high" } }
},
"usage": { "prompt_tokens": 306, "completion_tokens": 18, "total_tokens": 324 },
"extra_fields": { "provider": "openai" }
}Benchmarking
{
"_agreement_fail_threshold": 0.95,
"scenario": {
"state": "Customer support ticket #88231. Customer is on the Pro plan and has been a paying subscriber for 3 years. They write: 'This is my FIFTH email in two weeks and NOBODY has replied. You charged my card THREE times for a single order (#A-4021) at $149 each, so $447 total instead of $149. I have screenshots of all three charges. I have asked for a refund every single time and been completely ignored. If the two extra charges are not refunded TODAY I am filing a chargeback with my bank, cancelling my Pro subscription immediately, posting the screenshots on Twitter, and reporting you to the Better Business Bureau. Your competitor Acme offers the same features for less and I am seriously considering switching. This is completely unacceptable and I want a real person to fix this now.'",
"anchors": {
"is_frustrated": 1,
"is_angry": 1,
"mentions_refund": 1,
"threatens_chargeback": 1,
"threatens_cancellation": 1,
"mentions_competitor": 1,
"is_satisfied": 0,
"category": "billing",
"sentiment": "negative",
"customer_tier": "pro"
},
"questions": {
"is_frustrated": {"kind": "noul", "instructions": "Is the customer frustrated?"},
"is_angry": {"kind": "noul", "instructions": "Is the customer angry?"},
"is_satisfied": {"kind": "noul", "instructions": "Is the customer satisfied with the service?"},
"is_polite": {"kind": "noul", "instructions": "Is the message written in a calm, polite tone?"},
"mentions_refund": {"kind": "noul", "instructions": "Does the customer ask for a refund?"},
"threatens_chargeback": {"kind": "noul", "instructions": "Does the customer threaten a bank chargeback?"},
"threatens_cancellation": {"kind": "noul", "instructions": "Does the customer threaten to cancel their subscription?"},
"threatens_public_complaint": {"kind": "noul", "instructions": "Does the customer threaten to complain publicly (social media, review sites, or a regulator)?"},
"mentions_competitor": {"kind": "noul", "instructions": "Does the customer mention a competitor?"},
"is_billing_related": {"kind": "noul", "instructions": "Is the issue about billing or charges?"},
"is_repeat_contact": {"kind": "noul", "instructions": "Has the customer contacted support about this before?"},
"has_attached_evidence": {"kind": "noul", "instructions": "Does the customer say they have evidence such as screenshots?"},
"mentions_specific_amount": {"kind": "noul", "instructions": "Does the customer cite a specific dollar amount?"},
"is_long_time_customer": {"kind": "noul", "instructions": "Is this a long-time customer (multiple years)?"},
"requests_human_response": {"kind": "noul", "instructions": "Does the customer explicitly ask to speak to a real person?"},
"is_first_time_issue": {"kind": "noul", "instructions": "Is this the first time the customer is raising this issue?"},
"willing_to_wait": {"kind": "noul", "instructions": "Is the customer willing to wait patiently for a resolution?"},
"expresses_urgency": {"kind": "noul", "instructions": "Does the customer express that this is urgent?"},
"mentions_data_loss": {"kind": "noul", "instructions": "Does the customer report lost data?"},
"reports_product_crash": {"kind": "noul", "instructions": "Does the customer report the product crashing or a software defect?"},
"category": {"kind": "choice", "instructions": "Pick the ticket category", "criteria": {"billing": "charges, refunds, invoices", "bug": "product defects or crashes", "account": "login or account access", "other": "anything else"}},
"sentiment": {"kind": "choice", "instructions": "Overall sentiment of the message", "criteria": {"positive": "happy", "neutral": "matter of fact", "negative": "unhappy or hostile"}},
"primary_emotion": {"kind": "choice", "instructions": "The dominant emotion", "criteria": {"anger": "angry or hostile", "confusion": "confused", "sadness": "sad or disappointed", "joy": "happy"}},
"customer_tier": {"kind": "choice", "instructions": "Which plan is the customer on?", "criteria": {"free": "free plan", "pro": "pro plan", "enterprise": "enterprise plan"}},
"recommended_action": {"kind": "choice", "instructions": "The single best next action", "criteria": {"issue_refund": "refund the extra charges", "request_more_info": "ask for more details", "close_ticket": "close with no action", "ignore": "take no action"}},
"contact_channel": {"kind": "choice", "instructions": "Which channel did the customer use?", "criteria": {"email": "email or support ticket", "phone": "phone call", "chat": "live chat", "social": "social media"}},
"urgency": {"kind": "score", "instructions": "How urgently does this need a human reply?", "criteria": ["can wait a week", "should be answered this week", "should be answered today", "needs an immediate reply"]},
"severity": {"kind": "score", "instructions": "How severe is the customer's problem?", "criteria": ["minor", "moderate", "major", "severe"]},
"churn_risk": {"kind": "score", "instructions": "How likely is this customer to cancel?", "criteria": ["not at all likely", "somewhat likely", "very likely", "almost certain"]},
"customer_satisfaction": {"kind": "score", "instructions": "How satisfied is the customer right now?", "criteria": ["very unsatisfied", "unsatisfied", "neutral", "satisfied", "very satisfied"]}
}
}
}
Here are the results
| Combination | Score | Latency (ms) | Tokens | Status |
|---|---|---|---|---|
| typesafe/jev-1.13.0 | 30/30 | 920 | 2144 | REF |
| openai/gpt-5.6-luna | 30/30 | 6763 | 3534 | PASS |
| openai/gpt-5.6-terra | 30/30 | 11452 | 3596 | PASS |
| openai/gpt-5 | 30/30 | 25795 | 5901 | PASS |
| openai/gpt-5-mini | 29/30 | 36482 | 6487 | PASS |
| openai/gpt-4.1 | 30/30 | 3963 | 3336 | PASS |
| openai/gpt-4.1-mini | 29/30 | 13113 | 3365 | PASS |
| openai/gpt-4o | 30/30 | 7630 | 3470 | PASS |
| openai/gpt-4o-mini | 30/30 | 5852 | 3292 | PASS |
| anthropic/claude-opus-5 | 30/30 | 15667 | 8203 | PASS |
| anthropic/claude-sonnet-5 | 30/30 | 11768 | 8172 | PASS |
| anthropic/claude-fable-5-1 | 30/30 | 17789 | 8095 | PASS |
| anthropic/claude-haiku-4-5-20251001 | 30/30 | 7842 | 6784 | PASS |
| anthropic/claude-sonnet-4-5 | 30/30 | 16429 | 6784 | PASS |
| xai/grok-4.5 | 29/30 | 24335 | 6618 | PASS |
| xai/grok-4 | 30/30 | 11472 | 5958 | PASS |
| xai/grok-3 | 30/30 | 13600 | 5976 | PASS |
| xai/grok-3-mini | 29/30 | 12291 | 6147 | PASS |
| groq/openai/gpt-oss-120b | 30/30 | 4208 | 5194 | PASS |
| deepseek/deepseek-v4-pro | 30/30 | 10299 | 5841 | PASS |
| cohere/command-a-03-2025 | 30/30 | 9047 | 6487 | PASS |
| vertex/gemini-2.5-flash | 30/30 | 11127 | 1513 | PASS |
| vertex/claude-sonnet-4-6 | 30/30 | 16266 | 6477 | PASS |
| vertex/claude-opus-4-8 | 30/30 | 16091 | 8207 | PASS |
| gemini/gemini-2.5-flash | 30/30 | 9078 | 5307 | PASS |
| gemini/gemini-2.5-pro | 30/30 | 14811 | 5505 | PASS |
Try it
- Decisions quickstart - curl to first judgment in a minute
- API reference - full /v1/decisions and /typesafe/* schemas
- TypeSafe provider guide - complete field mapping to the native API
- TypeSafe - the jev models and System One