クイックスタート
Ollaya はオープンな意思決定モデルを自分のマシン上で実行します。モデルに状態(メッセージ、メール、チケット、JSON オブジェクト)といくつかの型付きの質問を渡すと、1 回の順伝播で、較正された確率を伴う型付きの答えを返します。
1. インストール
Linux と macOS の場合:
curl -fsSL https://ollaya.dev/install.sh | sh
Windows では PowerShell で:
irm https://ollaya.dev/install.ps1 | iex
またはデスクトップアプリを入手してください。同じコマンドラインとサーバーが同梱されています。
スクリプトは GitHub から最新リリースをダウンロードし、その sha256 を検証します。Linux と Windows では、NVIDIA GPU を見つけると CUDA ライブラリも取得します。Linux では、systemd が動作していて root 権限がある場合、127.0.0.1:11435 で API を提供するサービスを設定します。要件と Docker イメージはダウンロードを参照してください。
2. モデルを実行する
ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent refund ███████████████░ 0.91
is_urgent yes ███████████████░ 0.92
frustration 2.89 / 3 very angry or using stron… ██████████████░░ 0.86
refund_requested yes ████████████████ 0.99
churn_risk yes ████████████████ 0.99
ollaya run は、サーバーが起動していなければ起動し、初回使用時にモデルを取得して読み込みます。winnow:e4b は推奨モデルです。型付き意思決定で 0.722 の精度で、TypeSafe の Jev(0.738)に近く、RTX 4090 でこの 5 つの質問に 89 ms かかります。ダウンロードは 8 GB です。
NVIDIA GPU がありませんか?CPU で 1 秒の何分の一かで答える laya から始めてください。これは router です。英語のテキストを laya:en に、その他の言語(たとえばトルコ語)を laya:multilingual に送り、laya を取得すると両方が取得されます。Modelsが全モデルの精度と速度を比較します。
--preset NAMEは組み込みの質問セット(triage、email、guard、moderation、router、agent)を使います。指定しない場合、自前の質問を持たないモデルはtriageに答えます。--verboseは各選択肢の確率、ルーティングの判断、所要時間を追加で表示します。--format jsonは API の完全なレスポンスを出力します。- 状態がない場合、
ollaya runはパイプされた stdin を読むか、ターミナルでプロンプトを開きます。
3. 自分の質問を書く
質問をファイルに書きます:
{
"topic": {
"type": "choice",
"instructions": "What is this message about?",
"criteria": {
"billing": "Payments, invoices and refunds",
"access": "Login, passwords and permissions",
"other": "Anything else"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
}
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."
またはファイルを省略し、curl -d と同じように JSON をインラインで渡します({ で始まる値はファイルパスではなく JSON です):
ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."
または同じ質問を API に送ります:
curl http://localhost:11435/api/decide -d '{
"model": "winnow:e4b",
"state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
"questions": {
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
すべての答えは型付きで返ります。choice は選択肢ごとの確率を、score は期待される段階を、noul はその命題が成り立つ確率を返します。API リファレンスを参照してください。
4. 既存の TypeSafe クライアントを使う
Ollaya は TypeSafe の API も提供します。公式の TypeSafe Python SDK は変更なしで動作します:
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
TypeSafe 互換性を参照してください。
5. 質問をモデルに焼き込む
Modelfileは質問セットを、名前で実行できるモデルに変えます:
FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."