文件導航

安裝、下載並做出型別化決策

Laya-CoreML 在 Apple Silicon 上執行,要求 macOS 15+ 和 Python 3.11–3.13。本地釋出檢查使用 M3 Max / macOS 27.2。更早的 macOS 版本與 iOS 部署在這裡尚未測試。匯出的 ML Program 面向 macOS 15 / iOS 18。

python -m pip install laya-coreml
# Terminal demo and recording renderer:
python -m pip install 'laya-coreml[demo]'

推理會安裝 Core ML Tools、NumPy、Tokenizers、Safetensors 和 Hugging Face Hub。它不需要 PyTorch、Transformers 或 MLX。可選的 [convert] extra 會安裝 PyTorch,用於匯出原始 checkpoint。

選擇模型

aac6fef/ 下的 Hugging Face 包 預設裝置 總輸入容量 批 / 選項 預期用途
laya-coreml CPU + GPU 512 token 1 / 32 英文 Laya,421M
laya-multilingual-coreml CPU + GPU 1024 token 1 / 32 通用多語言,322M
laya-typed-decisions-coreml CPU + GPU 1024 token 1 / 32 上游 typed-decisions checkpoint
laya-multilingual-coreml-snake CPU + GPU 64 token 3 / 4 批處理的緊湊 Snake prompt
laya-multilingual-coreml-ane CPU + ANE 96 token 1 / 32 簡短決策,FP16 主體
laya-multilingual-coreml-ane-w8 CPU + ANE 96 token 1 / 32 近似 W8 調色盤權重,FP16 計算

ANE 包自帶其主機端 embedding/action 權重,以及未改動的 Core ML 主體。它們不需要原始的 checkpoint 目錄。96 token 的容量包含問題、選項描述、特殊標記和狀態。這些短匯出會拒絕放不下的請求。通用型模型在其完整 checkpoint 上下文上限處沿用上游的狀態截斷;選項描述也使用原始的問題字首預算。

Python API

import laya_coreml as laya

agent = laya.load("aac6fef/laya-multilingual-coreml")
result = agent.predict(
    "The customer asks for a refund of a duplicate payment.",
    {
        "department": {
            "type": "choice",
            "instructions": "Which department should handle this request?",
            "criteria": {
                "billing": "Payments, invoices, refunds, and duplicate charges.",
                "technical": "Broken features, errors, and product troubleshooting.",
                "sales": "Pricing, upgrades, and new purchases.",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is the request?",
            "criteria": ["low", "medium", "high"],
        },
        "refund": {
            "type": "noul",
            "instructions": "Does the customer request a refund?",
        },
    },
)
print(result["answers"])
print(result["usage"])  # output_tokens is always 0

choice 返回一個被選中的標籤,以及每個標籤的機率。score 返回期望的、從零開始的類別索引,它的量表和機率。noul 返回 true 的機率。答案還包含上游的置信度和 action-head 機率欄位。這些估計可能是錯的;本庫的驗證衡量的是轉換保真度,而非應用準確率。

predict 和 system_one 是別名。字典形式的 state 會按原始輸入約定序列化。問題按插入順序處理。大多數匯出使用批大小一;因此多個問題需要多次模型呼叫。Snake GPU 匯出最多批處理三個。雙向編碼器不會在不同問題之間快取上下文狀態。

下載一次,之後離線工作

hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/ane
agent = laya.load("./models/ane", local_files_only=True)
# Or use a previously downloaded shared Hub cache:
agent = laya.load("aac6fef/laya-multilingual-coreml-ane", local_files_only=True)

遠端 ID 會在初始化前下載,除非 local_files_only=True。之後的預測只使用本地的陣列和檔案。要復現一個精確的遠端快照,請傳入 revision="<Hub commit SHA>";釋出 commit ID 在 RELEASE.md。終端遊戲始終使用本地/快取的權重,缺失時會帶著一條下載命令報錯。

Hugging Face 的共享快取把模型檔案存為符號連結。在測試過的 macOS 版本上,Core ML 會把這些連結拷進它臨時編譯的模型,從而丟失權重檔案。載入器會自動把 Core ML 包本身物化為 ~/.cache/laya-coreml/packages/ 下的普通檔案,驗證其內容雜湊並複用該副本。設定 LAYA_COREML_CACHE 可更改這個快取根目錄。這隻多佔磁碟空間,不會多下載模型。用 hf download --local-dir 建立的目錄已經是普通檔案,無需複製。

載入器會識別包格式並選擇其預設計算單元。可用 compute_units="cpu_gpu"、"cpu_ne"、"cpu" 或 "all" 覆蓋,以做顯式實驗。cpu_ne 允許 CPU 工作和 ANE 工作;它不保證每個運算元都在 ANE 上執行。為普通 SDPA 匯出選擇它,並不會把那個匯出變成專用的 ANE 圖。

CLI

把問題字典儲存為 questions.json,然後執行:

laya-coreml predict ./models/ane --offline \
  --state 'The customer requests a refund.' --questions questions.json

laya-coreml predict 接受一個本地目錄或一個 Hub ID、可選的 --revision 和 --compute-units。--offline 阻止訪問 Hub。

從原始碼轉換

對於普通的 Core ML 路徑:

pip install 'laya-coreml[convert]'
laya-coreml convert laya-multilingual models/custom-multilingual

ANE 研究用的匯出器在 Git checkout 裡,不在推理 wheel 中:

git clone https://github.com/mizorewww/laya-coreml
cd laya-coreml
pip install -e '.[convert,dev,research]'
python -m experiments.ane_engineering.probe --source laya-multilingual \
  --kind body --length 96 --output models/ane96

對應的工作流見 轉換髮現、ANE 工程 和 Snake 控制與錄製。