快速開始
Ollaya 把開放式決策模型跑在你自己的機器上。你給模型一個狀態(一條訊息、一封郵件、一張工單、一個 JSON 物件),再加幾個型別化的問題,它在一次前向傳播中返回帶校準機率的型別化答案。
1. 安裝
在 Linux 和 macOS 上:
curl -fsSL https://ollaya.dev/install.sh | sh
在 Windows 上,用 PowerShell:
irm https://ollaya.dev/install.ps1 | iex
或者裝桌面應用,它把同一個命令列和伺服器一起打包。
指令碼從 GitHub 下載最新發行版並校驗它的 sha256。在 Linux 和 Windows 上,它發現 NVIDIA GPU 時還會拉取 CUDA 庫。在 Linux 上,有 systemd 且有 root 許可權時,它會裝一個服務,在 127.0.0.1:11435 上提供 API。需求和 Docker 映象見下載。
2. 執行一個模型
ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent refund ███████████████░ 0.91
is_urgent yes ███████████████░ 0.92
frustration 2.89 / 3 very angry or using stron… ██████████████░░ 0.86
refund_requested yes ████████████████ 0.99
churn_risk yes ████████████████ 0.99
ollaya run 會在伺服器沒執行時啟動它,首次使用時拉取模型並載入。winnow:e4b 是推薦的模型:型別化決策上 0.722 的準確率,接近 TypeSafe 的 Jev(0.738),在 RTX 4090 上回答這五個問題用 89 ms。它需要下載 8 GB。
沒有 NVIDIA GPU?從 laya 開始,它在 CPU 上零點幾秒就能回答。它是一個 router:把英語文本發給 laya:en,把其他語言(比如土耳其語)發給 laya:multilingual,而拉取 laya 會同時拉取兩者。Models 比較每個模型的準確率和速度。
--preset NAME使用一組內建問題:triage、email、guard、moderation、router或agent。不用它時,沒有自帶問題的模型回答triage。--verbose額外列印每個選項的機率、路由決策和各項耗時。--format json列印完整的 API 響應。- 沒有狀態時,
ollaya run讀取管道傳入的 stdin,或在終端上開啟一個提示介面。
3. 提出你自己的問題
把問題寫進一個檔案:
{
"topic": {
"type": "choice",
"instructions": "What is this message about?",
"criteria": {
"billing": "Payments, invoices and refunds",
"access": "Login, passwords and permissions",
"other": "Anything else"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
}
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."
或者跳過檔案,按 curl -d 的方式在命令列上直接傳 JSON(以 { 開頭的值是 JSON,不是檔案路徑):
ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."
或者把同樣的題目發給 API:
curl http://localhost:11435/api/decide -d '{
"model": "winnow:e4b",
"state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
"questions": {
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
每個答案都會帶型別地返回:一個 choice,每個選項各有一個機率;一個 score,是期望檔位;一個 noul,是這句話成立的機率。見 API 參考。
4. 使用現有的 TypeSafe 客戶端
Ollaya 也提供 TypeSafe 的 API。官方的 TypeSafe Python SDK 無需改動即可使用:
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
見 TypeSafe 相容性。
5. 把你的問題烤進模型
Modelfile 把一組問題變成一個可以用名字執行的模型:
FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."