文档导航

安装、下载并做出类型化决策

Laya-CoreML 在 Apple Silicon 上运行,要求 macOS 15+ 和 Python 3.11–3.13。本地发布检查使用 M3 Max / macOS 27.2。更早的 macOS 版本与 iOS 部署在这里尚未测试。导出的 ML Program 面向 macOS 15 / iOS 18。

python -m pip install laya-coreml
# Terminal demo and recording renderer:
python -m pip install 'laya-coreml[demo]'

推理会安装 Core ML Tools、NumPy、Tokenizers、Safetensors 和 Hugging Face Hub。它不需要 PyTorch、Transformers 或 MLX。可选的 [convert] extra 会安装 PyTorch,用于导出原始 checkpoint。

选择模型

aac6fef/ 下的 Hugging Face 包 默认设备 总输入容量 批 / 选项 预期用途
laya-coreml CPU + GPU 512 token 1 / 32 英文 Laya,421M
laya-multilingual-coreml CPU + GPU 1024 token 1 / 32 通用多语言,322M
laya-typed-decisions-coreml CPU + GPU 1024 token 1 / 32 上游 typed-decisions checkpoint
laya-multilingual-coreml-snake CPU + GPU 64 token 3 / 4 批处理的紧凑 Snake prompt
laya-multilingual-coreml-ane CPU + ANE 96 token 1 / 32 简短决策,FP16 主体
laya-multilingual-coreml-ane-w8 CPU + ANE 96 token 1 / 32 近似 W8 调色板权重,FP16 计算

ANE 包自带其主机端 embedding/action 权重,以及未改动的 Core ML 主体。它们不需要原始的 checkpoint 目录。96 token 的容量包含问题、选项描述、特殊标记和状态。这些短导出会拒绝放不下的请求。通用型模型在其完整 checkpoint 上下文上限处沿用上游的状态截断;选项描述也使用原始的问题前缀预算。

Python API

import laya_coreml as laya

agent = laya.load("aac6fef/laya-multilingual-coreml")
result = agent.predict(
    "The customer asks for a refund of a duplicate payment.",
    {
        "department": {
            "type": "choice",
            "instructions": "Which department should handle this request?",
            "criteria": {
                "billing": "Payments, invoices, refunds, and duplicate charges.",
                "technical": "Broken features, errors, and product troubleshooting.",
                "sales": "Pricing, upgrades, and new purchases.",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is the request?",
            "criteria": ["low", "medium", "high"],
        },
        "refund": {
            "type": "noul",
            "instructions": "Does the customer request a refund?",
        },
    },
)
print(result["answers"])
print(result["usage"])  # output_tokens is always 0

choice 返回一个被选中的标签,以及每个标签的概率。score 返回期望的、从零开始的类别索引,它的量表和概率。noul 返回 true 的概率。答案还包含上游的置信度和 action-head 概率字段。这些估计可能是错的;本库的验证衡量的是转换保真度,而非应用准确率。

predict 和 system_one 是别名。字典形式的 state 会按原始输入约定序列化。问题按插入顺序处理。大多数导出使用批大小一;因此多个问题需要多次模型调用。Snake GPU 导出最多批处理三个。双向编码器不会在不同问题之间缓存上下文状态。

下载一次,之后离线工作

hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/ane
agent = laya.load("./models/ane", local_files_only=True)
# Or use a previously downloaded shared Hub cache:
agent = laya.load("aac6fef/laya-multilingual-coreml-ane", local_files_only=True)

远程 ID 会在初始化前下载,除非 local_files_only=True。之后的预测只使用本地的数组和文件。要复现一个精确的远程快照,请传入 revision="<Hub commit SHA>";发布 commit ID 在 RELEASE.md。终端游戏始终使用本地/缓存的权重,缺失时会带着一条下载命令报错。

Hugging Face 的共享缓存把模型文件存为符号链接。在测试过的 macOS 版本上,Core ML 会把这些链接拷进它临时编译的模型,从而丢失权重文件。加载器会自动把 Core ML 包本身物化为 ~/.cache/laya-coreml/packages/ 下的普通文件,验证其内容哈希并复用该副本。设置 LAYA_COREML_CACHE 可更改这个缓存根目录。这只多占磁盘空间,不会多下载模型。用 hf download --local-dir 创建的目录已经是普通文件,无需拷贝。

加载器会识别包格式并选择其默认计算单元。可用 compute_units="cpu_gpu"、"cpu_ne"、"cpu" 或 "all" 覆盖,以做显式实验。cpu_ne 允许 CPU 工作和 ANE 工作;它不保证每个算子都在 ANE 上执行。为普通 SDPA 导出选择它,并不会把那个导出变成专用的 ANE 图。

CLI

把问题字典保存为 questions.json,然后运行:

laya-coreml predict ./models/ane --offline \
  --state 'The customer requests a refund.' --questions questions.json

laya-coreml predict 接受一个本地目录或一个 Hub ID、可选的 --revision 和 --compute-units。--offline 阻止访问 Hub。

从源码转换

对于普通的 Core ML 路径:

pip install 'laya-coreml[convert]'
laya-coreml convert laya-multilingual models/custom-multilingual

ANE 研究用的导出器在 Git checkout 里,不在推理 wheel 中:

git clone https://github.com/mizorewww/laya-coreml
cd laya-coreml
pip install -e '.[convert,dev,research]'
python -m experiments.ane_engineering.probe --source laya-multilingual \
  --kind body --length 96 --output models/ane96

对应的工作流见 转换发现、ANE 工程 和 Snake 控制与录制。