快速开始
Ollaya 把开放式决策模型跑在你自己的机器上。你给模型一个状态(一条消息、一封邮件、一张工单、一个 JSON 对象),再加几个类型化的问题,它在一次前向传播中返回带校准概率的类型化答案。
1. 安装
在 Linux 和 macOS 上:
curl -fsSL https://ollaya.dev/install.sh | sh
在 Windows 上,用 PowerShell:
irm https://ollaya.dev/install.ps1 | iex
或者装桌面应用,它把同一个命令行和服务器一起打包。
脚本从 GitHub 下载最新发行版并校验它的 sha256。在 Linux 和 Windows 上,它发现 NVIDIA GPU 时还会拉取 CUDA 库。在 Linux 上,有 systemd 且有 root 权限时,它会装一个服务,在 127.0.0.1:11435 上提供 API。需求和 Docker 镜像见下载。
2. 运行一个模型
ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent refund ███████████████░ 0.91
is_urgent yes ███████████████░ 0.92
frustration 2.89 / 3 very angry or using stron… ██████████████░░ 0.86
refund_requested yes ████████████████ 0.99
churn_risk yes ████████████████ 0.99
ollaya run 会在服务器没运行时启动它,首次使用时拉取模型并加载。winnow:e4b 是推荐的模型:类型化决策上 0.722 的准确率,接近 TypeSafe 的 Jev(0.738),在 RTX 4090 上回答这五个问题用 89 ms。它需要下载 8 GB。
没有 NVIDIA GPU?从 laya 开始,它在 CPU 上零点几秒就能回答。它是一个 router:把英语文本发给 laya:en,把其他语言(比如土耳其语)发给 laya:multilingual,而拉取 laya 会同时拉取两者。Models 比较每个模型的准确率和速度。
--preset NAME使用一组内置问题:triage、email、guard、moderation、router或agent。不用它时,没有自带问题的模型回答triage。--verbose额外打印每个选项的概率、路由决策和各项耗时。--format json打印完整的 API 响应。- 没有状态时,
ollaya run读取管道传入的 stdin,或在终端上打开一个提示界面。
3. 提出你自己的问题
把问题写进一个文件:
{
"topic": {
"type": "choice",
"instructions": "What is this message about?",
"criteria": {
"billing": "Payments, invoices and refunds",
"access": "Login, passwords and permissions",
"other": "Anything else"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
}
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."
或者跳过文件,按 curl -d 的方式在命令行上直接传 JSON(以 { 开头的值是 JSON,不是文件路径):
ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."
或者把同样的题目发给 API:
curl http://localhost:11435/api/decide -d '{
"model": "winnow:e4b",
"state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
"questions": {
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
每个答案都会带类型地返回:一个 choice,每个选项各有一个概率;一个 score,是期望档位;一个 noul,是这句话成立的概率。见 API 参考。
4. 使用现有的 TypeSafe 客户端
Ollaya 也提供 TypeSafe 的 API。官方的 TypeSafe Python SDK 无需改动即可使用:
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
见 TypeSafe 兼容性。
5. 把你的问题烤进模型
Modelfile 把一组问题变成一个可以用名字运行的模型:
FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."