ドキュメント

クイックスタート

Ollaya はオープンな意思決定モデルを自分のマシン上で実行します。モデルに状態(メッセージ、メール、チケット、JSON オブジェクト)といくつかの型付きの質問を渡すと、1 回の順伝播で、較正された確率を伴う型付きの答えを返します。

1. インストール

Linux と macOS の場合:

curl -fsSL https://ollaya.dev/install.sh | sh

Windows では PowerShell で:

irm https://ollaya.dev/install.ps1 | iex

またはデスクトップアプリを入手してください。同じコマンドラインとサーバーが同梱されています。

スクリプトは GitHub から最新リリースをダウンロードし、その sha256 を検証します。Linux と Windows では、NVIDIA GPU を見つけると CUDA ライブラリも取得します。Linux では、systemd が動作していて root 権限がある場合、127.0.0.1:11435 で API を提供するサービスを設定します。要件と Docker イメージはダウンロードを参照してください。

2. モデルを実行する

ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent            refund                                ███████████████░ 0.91
is_urgent         yes                                   ███████████████░ 0.92
frustration       2.89 / 3  very angry or using stron…  ██████████████░░ 0.86
refund_requested  yes                                   ████████████████ 0.99
churn_risk        yes                                   ████████████████ 0.99

ollaya run は、サーバーが起動していなければ起動し、初回使用時にモデルを取得して読み込みます。winnow:e4b は推奨モデルです。型付き意思決定で 0.722 の精度で、TypeSafe の Jev(0.738)に近く、RTX 4090 でこの 5 つの質問に 89 ms かかります。ダウンロードは 8 GB です。

NVIDIA GPU がありませんか?CPU で 1 秒の何分の一かで答える laya から始めてください。これは router です。英語のテキストを laya:en に、その他の言語(たとえばトルコ語)を laya:multilingual に送り、laya を取得すると両方が取得されます。Modelsが全モデルの精度と速度を比較します。

  • --preset NAME は組み込みの質問セット(triage、email、guard、moderation、router、agent)を使います。指定しない場合、自前の質問を持たないモデルは triage に答えます。
  • --verbose は各選択肢の確率、ルーティングの判断、所要時間を追加で表示します。
  • --format json は API の完全なレスポンスを出力します。
  • 状態がない場合、ollaya run はパイプされた stdin を読むか、ターミナルでプロンプトを開きます。

3. 自分の質問を書く

質問をファイルに書きます:

{
  "topic": {
    "type": "choice",
    "instructions": "What is this message about?",
    "criteria": {
      "billing": "Payments, invoices and refunds",
      "access": "Login, passwords and permissions",
      "other": "Anything else"
    }
  },
  "urgency": {
    "type": "score",
    "instructions": "How urgent is this?",
    "criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
  },
  "angry": {
    "type": "noul",
    "instructions": "Is the customer angry?"
  }
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."

またはファイルを省略し、curl -d と同じように JSON をインラインで渡します({ で始まる値はファイルパスではなく JSON です):

ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."

または同じ質問を API に送ります:

curl http://localhost:11435/api/decide -d '{
  "model": "winnow:e4b",
  "state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
  "questions": {
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

すべての答えは型付きで返ります。choice は選択肢ごとの確率を、score は期待される段階を、noul はその命題が成り立つ確率を返します。API リファレンスを参照してください。

4. 既存の TypeSafe クライアントを使う

Ollaya は TypeSafe の API も提供します。公式の TypeSafe Python SDK は変更なしで動作します:

export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b

TypeSafe 互換性を参照してください。

5. 質問をモデルに焼き込む

Modelfileは質問セットを、名前で実行できるモデルに変えます:

FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."

次のステップ