文件導航

快速開始

Ollaya 把開放式決策模型跑在你自己的機器上。你給模型一個狀態(一條訊息、一封郵件、一張工單、一個 JSON 物件),再加幾個型別化的問題,它在一次前向傳播中返回帶校準機率的型別化答案。

1. 安裝

在 Linux 和 macOS 上:

curl -fsSL https://ollaya.dev/install.sh | sh

在 Windows 上,用 PowerShell:

irm https://ollaya.dev/install.ps1 | iex

或者裝桌面應用,它把同一個命令列和伺服器一起打包。

指令碼從 GitHub 下載最新發行版並校驗它的 sha256。在 Linux 和 Windows 上,它發現 NVIDIA GPU 時還會拉取 CUDA 庫。在 Linux 上,有 systemd 且有 root 許可權時,它會裝一個服務,在 127.0.0.1:11435 上提供 API。需求和 Docker 映象見下載。

2. 執行一個模型

ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent            refund                                ███████████████░ 0.91
is_urgent         yes                                   ███████████████░ 0.92
frustration       2.89 / 3  very angry or using stron…  ██████████████░░ 0.86
refund_requested  yes                                   ████████████████ 0.99
churn_risk        yes                                   ████████████████ 0.99

ollaya run 會在伺服器沒執行時啟動它,首次使用時拉取模型並載入。winnow:e4b 是推薦的模型:型別化決策上 0.722 的準確率,接近 TypeSafe 的 Jev(0.738),在 RTX 4090 上回答這五個問題用 89 ms。它需要下載 8 GB。

沒有 NVIDIA GPU?從 laya 開始,它在 CPU 上零點幾秒就能回答。它是一個 router:把英語文本發給 laya:en,把其他語言(比如土耳其語)發給 laya:multilingual,而拉取 laya 會同時拉取兩者。Models 比較每個模型的準確率和速度。

  • --preset NAME 使用一組內建問題:triage、email、guard、moderation、router 或 agent。不用它時,沒有自帶問題的模型回答 triage。
  • --verbose 額外列印每個選項的機率、路由決策和各項耗時。
  • --format json 列印完整的 API 響應。
  • 沒有狀態時,ollaya run 讀取管道傳入的 stdin,或在終端上開啟一個提示介面。

3. 提出你自己的問題

把問題寫進一個檔案:

{
  "topic": {
    "type": "choice",
    "instructions": "What is this message about?",
    "criteria": {
      "billing": "Payments, invoices and refunds",
      "access": "Login, passwords and permissions",
      "other": "Anything else"
    }
  },
  "urgency": {
    "type": "score",
    "instructions": "How urgent is this?",
    "criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
  },
  "angry": {
    "type": "noul",
    "instructions": "Is the customer angry?"
  }
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."

或者跳過檔案,按 curl -d 的方式在命令列上直接傳 JSON(以 { 開頭的值是 JSON,不是檔案路徑):

ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."

或者把同樣的題目發給 API:

curl http://localhost:11435/api/decide -d '{
  "model": "winnow:e4b",
  "state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
  "questions": {
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

每個答案都會帶型別地返回:一個 choice,每個選項各有一個機率;一個 score,是期望檔位;一個 noul,是這句話成立的機率。見 API 參考。

4. 使用現有的 TypeSafe 客戶端

Ollaya 也提供 TypeSafe 的 API。官方的 TypeSafe Python SDK 無需改動即可使用:

export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b

見 TypeSafe 相容性。

5. 把你的問題烤進模型

Modelfile 把一組問題變成一個可以用名字執行的模型:

FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."

下一步