Jev は TypeSafe AI の最初のSystem One model。
構造化されていないプログラムの状態(state)と型付きの質問を受け取り、型付きの判断を確率・confidenceとともに 1回の並列パスで 返す。
特徴¶
新しいモデルアーキテクチャ + 並列サンプラー:自己回帰でトークンを逐次生成するのではなく、全質問の答えを1回のクエリで同時に計算する
型安全:出力は事前に定義された選択肢のどれかなので、型エラーが0%(スキーマで保証される)
キャリブレーション:RLCDで学習され、確信度が高いほど正解率も高い
2年間のステルス期間を経て、新しいアーキテクチャと学習方法で開発された
アーキテクチャや学習の詳細は公開されていない。
性能(公称値)¶
| 項目 | 値 |
|---|---|
| レイテンシ | エンドツーエンドで 70〜500ms |
| 速度 | 同等の性能のフロンティアLLMより 40〜200倍速い |
| 入力トークン価格 | $0.042 / 100万トークン |
| 出力トークン価格 | 無料(too cheap to meter) |
| 型エラー | 0% |
ワークフロー評価では、GPT-6 Astra や Fable 5.1 と比較して「約2桁にわたってパレートフロンティアを押さえている」としている。
注意点(TypeSafe自身が挙げているもの)¶
汎用的な文字列生成の能力は捨てている
Choiceの選択肢数は最大255
入力はデータ構造+テキストのみで、画像・音声・動画は未対応
レイテンシは米国西海岸のサーバーから計測
評価ワークフローは自社のcapabilitiesチームが作ったもので、バイアスがありうる。数字は「実世界での効果の上の方」とのこと
また、Jev 1.13 Jaggedness として既知の苦手領域が公開されている。
使い方¶
HTTP API¶
エンドポイントは POST https://api.typesafe.ai/v1/systemone。デフォルトのモデルは jev-latest。
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": {
"type": "noul",
"instructions": "Does this message express urgency?"
}
}
}
EOFPython SDK¶
pip install typesafe-sdk # Python >= 3.10from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # 環境変数 TYPESAFE_API_KEY を使う
response = client.system_one(
state=state,
questions={
"refund_requested": Noul(
instructions="Does `ticket_message` request a refund?",
),
"request_type": Choice(
instructions="What is the main request in `ticket_message`?",
criteria={
"refund": "Customer wants money returned.",
"rebooking": "Customer wants replacement flight.",
"information": "Customer asking for info only.",
},
),
"frustration": Score(
instructions="How frustrated in `ticket_message`?",
criteria=[
"Calm and neutral.",
"Concerned but civil.",
"Very angry or strong language.",
],
),
},
)
response.answers["refund_requested"].noul # 0〜1の確率
response.answers["request_type"].choice # "refund" など
response.answers["frustration"].score # 0〜2Choiceの
criteriaは dict(選択肢名 → 説明)Scoreの
criteriaは順序付きのlist(低い → 高い)非同期版の
AsyncTypeSafeClientもあるJavaScript/TypeScript SDK もある
他にも、Playground やClaude Code用のAgent Skill(claude plugin install typesafe@typesafe-ai)が用意されている。
confidenceで分岐する例¶
response = client.system_one(
state=user_message,
questions={
"action": Choice(
instructions="What is the user trying to do?",
criteria={
"check_balance": "View account balance",
"approve_transfer": "Approve withdrawal",
"support": "Get help",
},
),
},
)
action = response.answers["action"]
if action.confidence < 0.5:
route_to_human(user_message)
elif action.choice == "approve_transfer" and action.confidence > 0.9:
confirm_then_execute(account_id)
else:
ask_user_to_confirm(account_id)ユースケース¶
公式のCookbooksで紹介されている例:
問い合わせの振り分け(Intent Routing)
RAGで取得したパッセージの関連性フィルタリング、リランキング
引用の正しさのチェック、LLMの入出力のガードレール
自然言語のリクエストから型付き関数の呼び出しへの変換(Function Calling)
182件のスキルカタログからのスキル選択
エンティティアライメント(商品ペアのマッチング)
日付抽出、正規表現で取った候補の検証
深い階層を持つ分類(ビームサーチと組み合わせる)
LLMの「前段・後段の判断役」として組み合わせて使うパターンが多い。
system-one-adapter-python¶
TypeSafe公式の、system_one APIの ドロップイン代替。評価のバックエンドを TypeSafe の代わりに OpenAI / Anthropic / Gemini などのLLM API にする。
同じコードで「Jevを使った場合」と「LLMを使った場合」の コスト・レイテンシ・精度を比較する ためのツールという位置づけ。
pip install 'system-one-adapter[openai]' # OpenAI互換
pip install 'system-one-adapter[anthropic]' # Anthropic
pip install 'system-one-adapter[gemini]' # Geminifrom system_one_adapter import SystemOneAdapterClient, Noul
client = SystemOneAdapterClient(
structured_outputs=True, # ネイティブのstructured outputsを使う(Falseならプロンプトで JSON を指示)
llm_answer_mode="probabilities", # LLMに確率を答えさせる(離散値を答えさせるモードもある)
normalize_probabilities=True, # LLMが返した確率を正規化する
)
response = client.system_one(
state="This book was a delight to read.",
questions={"positive": Noul(instructions="The book review is positive.")},
provider="openai",
model="gpt-4o-mini",
)レスポンスには
input_tokens_total/output_tokens_total/ レイテンシなどの使用量、リトライ回数、全LLM呼び出しのデバッグ情報が入る壊れたJSONに対するリトライ回数や、リトライポリシーを設定できる
非同期版の
AsyncSystemOneAdapterClientもある
LLMの場合、確率は「LLMに数値を言わせたもの」なので、Jevのようなキャリブレーションは保証されない点に注意。