安装、下载并做出类型化决策
Laya-CoreML 在 Apple Silicon 上运行,要求 macOS 15+ 和 Python 3.11–3.13。本地发布检查使用 M3 Max / macOS 27.2。更早的 macOS 版本与 iOS 部署在这里尚未测试。导出的 ML Program 面向 macOS 15 / iOS 18。
python -m pip install laya-coreml
# Terminal demo and recording renderer:
python -m pip install 'laya-coreml[demo]'
推理会安装 Core ML Tools、NumPy、Tokenizers、Safetensors 和 Hugging Face Hub。它不需要 PyTorch、Transformers 或 MLX。可选的 [convert] extra 会安装 PyTorch,用于导出原始 checkpoint。
选择模型
aac6fef/ 下的 Hugging Face 包 |
默认设备 | 总输入容量 | 批 / 选项 | 预期用途 |
|---|---|---|---|---|
| laya-coreml | CPU + GPU | 512 token | 1 / 32 | 英文 Laya,421M |
| laya-multilingual-coreml | CPU + GPU | 1024 token | 1 / 32 | 通用多语言,322M |
| laya-typed-decisions-coreml | CPU + GPU | 1024 token | 1 / 32 | 上游 typed-decisions checkpoint |
| laya-multilingual-coreml-snake | CPU + GPU | 64 token | 3 / 4 | 批处理的紧凑 Snake prompt |
| laya-multilingual-coreml-ane | CPU + ANE | 96 token | 1 / 32 | 简短决策,FP16 主体 |
| laya-multilingual-coreml-ane-w8 | CPU + ANE | 96 token | 1 / 32 | 近似 W8 调色板权重,FP16 计算 |
ANE 包自带其主机端 embedding/action 权重,以及未改动的 Core ML 主体。它们不需要原始的 checkpoint 目录。96 token 的容量包含问题、选项描述、特殊标记和状态。这些短导出会拒绝放不下的请求。通用型模型在其完整 checkpoint 上下文上限处沿用上游的状态截断;选项描述也使用原始的问题前缀预算。
Python API
import laya_coreml as laya
agent = laya.load("aac6fef/laya-multilingual-coreml")
result = agent.predict(
"The customer asks for a refund of a duplicate payment.",
{
"department": {
"type": "choice",
"instructions": "Which department should handle this request?",
"criteria": {
"billing": "Payments, invoices, refunds, and duplicate charges.",
"technical": "Broken features, errors, and product troubleshooting.",
"sales": "Pricing, upgrades, and new purchases.",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is the request?",
"criteria": ["low", "medium", "high"],
},
"refund": {
"type": "noul",
"instructions": "Does the customer request a refund?",
},
},
)
print(result["answers"])
print(result["usage"]) # output_tokens is always 0
choice 返回一个被选中的标签,以及每个标签的概率。score 返回期望的、从零开始的类别索引,它的量表和概率。noul 返回 true 的概率。答案还包含上游的置信度和 action-head 概率字段。这些估计可能是错的;本库的验证衡量的是转换保真度,而非应用准确率。
predict 和 system_one 是别名。字典形式的 state 会按原始输入约定序列化。问题按插入顺序处理。大多数导出使用批大小一;因此多个问题需要多次模型调用。Snake GPU 导出最多批处理三个。双向编码器不会在不同问题之间缓存上下文状态。
下载一次,之后离线工作
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/ane
agent = laya.load("./models/ane", local_files_only=True)
# Or use a previously downloaded shared Hub cache:
agent = laya.load("aac6fef/laya-multilingual-coreml-ane", local_files_only=True)
远程 ID 会在初始化前下载,除非 local_files_only=True。之后的预测只使用本地的数组和文件。要复现一个精确的远程快照,请传入 revision="<Hub commit SHA>";发布 commit ID 在 RELEASE.md。终端游戏始终使用本地/缓存的权重,缺失时会带着一条下载命令报错。
Hugging Face 的共享缓存把模型文件存为符号链接。在测试过的 macOS 版本上,Core ML 会把这些链接拷进它临时编译的模型,从而丢失权重文件。加载器会自动把 Core ML 包本身物化为 ~/.cache/laya-coreml/packages/ 下的普通文件,验证其内容哈希并复用该副本。设置 LAYA_COREML_CACHE 可更改这个缓存根目录。这只多占磁盘空间,不会多下载模型。用 hf download --local-dir 创建的目录已经是普通文件,无需拷贝。
加载器会识别包格式并选择其默认计算单元。可用 compute_units="cpu_gpu"、"cpu_ne"、"cpu" 或 "all" 覆盖,以做显式实验。cpu_ne 允许 CPU 工作和 ANE 工作;它不保证每个算子都在 ANE 上执行。为普通 SDPA 导出选择它,并不会把那个导出变成专用的 ANE 图。
CLI
把问题字典保存为 questions.json,然后运行:
laya-coreml predict ./models/ane --offline \
--state 'The customer requests a refund.' --questions questions.json
laya-coreml predict 接受一个本地目录或一个 Hub ID、可选的 --revision 和 --compute-units。--offline 阻止访问 Hub。
从源码转换
对于普通的 Core ML 路径:
pip install 'laya-coreml[convert]'
laya-coreml convert laya-multilingual models/custom-multilingual
ANE 研究用的导出器在 Git checkout 里,不在推理 wheel 中:
git clone https://github.com/mizorewww/laya-coreml
cd laya-coreml
pip install -e '.[convert,dev,research]'
python -m experiments.ane_engineering.probe --source laya-multilingual \
--kind body --length 96 --output models/ane96
对应的工作流见 转换发现、ANE 工程 和 Snake 控制与录制。