文档导航

快速开始

Ollaya 把开放式决策模型跑在你自己的机器上。你给模型一个状态(一条消息、一封邮件、一张工单、一个 JSON 对象),再加几个类型化的问题,它在一次前向传播中返回带校准概率的类型化答案。

1. 安装

在 Linux 和 macOS 上:

curl -fsSL https://ollaya.dev/install.sh | sh

在 Windows 上,用 PowerShell:

irm https://ollaya.dev/install.ps1 | iex

或者装桌面应用,它把同一个命令行和服务器一起打包。

脚本从 GitHub 下载最新发行版并校验它的 sha256。在 Linux 和 Windows 上,它发现 NVIDIA GPU 时还会拉取 CUDA 库。在 Linux 上,有 systemd 且有 root 权限时,它会装一个服务,在 127.0.0.1:11435 上提供 API。需求和 Docker 镜像见下载。

2. 运行一个模型

ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent            refund                                ███████████████░ 0.91
is_urgent         yes                                   ███████████████░ 0.92
frustration       2.89 / 3  very angry or using stron…  ██████████████░░ 0.86
refund_requested  yes                                   ████████████████ 0.99
churn_risk        yes                                   ████████████████ 0.99

ollaya run 会在服务器没运行时启动它,首次使用时拉取模型并加载。winnow:e4b 是推荐的模型:类型化决策上 0.722 的准确率,接近 TypeSafe 的 Jev(0.738),在 RTX 4090 上回答这五个问题用 89 ms。它需要下载 8 GB。

没有 NVIDIA GPU?从 laya 开始,它在 CPU 上零点几秒就能回答。它是一个 router:把英语文本发给 laya:en,把其他语言(比如土耳其语)发给 laya:multilingual,而拉取 laya 会同时拉取两者。Models 比较每个模型的准确率和速度。

  • --preset NAME 使用一组内置问题:triage、email、guard、moderation、router 或 agent。不用它时,没有自带问题的模型回答 triage。
  • --verbose 额外打印每个选项的概率、路由决策和各项耗时。
  • --format json 打印完整的 API 响应。
  • 没有状态时,ollaya run 读取管道传入的 stdin,或在终端上打开一个提示界面。

3. 提出你自己的问题

把问题写进一个文件:

{
  "topic": {
    "type": "choice",
    "instructions": "What is this message about?",
    "criteria": {
      "billing": "Payments, invoices and refunds",
      "access": "Login, passwords and permissions",
      "other": "Anything else"
    }
  },
  "urgency": {
    "type": "score",
    "instructions": "How urgent is this?",
    "criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
  },
  "angry": {
    "type": "noul",
    "instructions": "Is the customer angry?"
  }
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."

或者跳过文件,按 curl -d 的方式在命令行上直接传 JSON(以 { 开头的值是 JSON,不是文件路径):

ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."

或者把同样的题目发给 API:

curl http://localhost:11435/api/decide -d '{
  "model": "winnow:e4b",
  "state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
  "questions": {
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

每个答案都会带类型地返回:一个 choice,每个选项各有一个概率;一个 score,是期望档位;一个 noul,是这句话成立的概率。见 API 参考。

4. 使用现有的 TypeSafe 客户端

Ollaya 也提供 TypeSafe 的 API。官方的 TypeSafe Python SDK 无需改动即可使用:

export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b

见 TypeSafe 兼容性。

5. 把你的问题烤进模型

Modelfile 把一组问题变成一个可以用名字运行的模型:

FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."

下一步