技能建議
在一輪 agent 互動中,從 Nous Research 的 Hermes 目錄的 182 個技能裡最多挑出一個,用兩次 TypeSafe 請求來排序並複查靠前的候選。
agent 通過截斷並把所有技能都塞進系統訊息來挑選技能,這會增加成本、拉低技能選擇的表現, 並讓會話後面的上下文腐化。我們用每輪兩次 TypeSafe 請求來應對:一次給技能排序, 一次驗證選中的那個,把錯誤的技能載入減少一半以上。
技能名冊很大的 agent,幾乎是在毫無資訊的情況下做選擇。名冊是以索引的形式
送到它手上的:每個技能一行,描述被截斷,免得全文
擠佔對話。這裡用的 agent harness Hermes 預設把它截到 60 個
字元。舉例來說,在這個寬度下,編輯 .pptx 檔案的技能
讀起來和創作這類檔案的技能幾乎一樣。讓 agent 做一份路演 deck,它可能
載入錯的那個。碰到沒有任何技能合適的一輪,它可能還是會
載入一個,因為一串名字本身就在引誘它去猜。
這個 cookbook 不動那些描述,改用漸進式披露: 先把 182 個技能都廉價地讀一遍,再詳讀其中三個。在「載入哪個技能(如果有的話)」 這個決定前面,放兩次 TypeSafe 請求。第一次拿名冊裡的每個技能 和使用者這一輪做比較排序,並判斷這一輪到底需不需要技能。 第二次只重讀前三名,這次帶上每個技能的完整描述和它指令的 開頭,並且允許把它們全部拒絕。
勝出的那個名字會變成 agent 那一輪系統提示詞裡多出來的一行:
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user
actually asked for.
</skill_relevance>
agent 保留完整的索引和自己的判斷,那一行只是告訴它先去查哪個
條目。名冊本身從不改變,所以任何針對它的字首快取依然
有效。在針對 claude-haiku-4-5-20251001 的 488 次請求中,使用 Hermes 名冊裡的技能:
| 載入錯誤的技能 | 沒有合適的也載入一個 | |
|---|---|---|
| 單獨的 agent,只有自己的名冊 | 16.8% | 9.8% |
| 帶 TypeSafe 建議的 agent | 7.3% | 4.0% |
| 直接拿到正確答案的 agent | 2.5% | 1.2% |
第三行說明,犯錯的下限不是零:即使把正確的技能直接交給 agent, 它也不總會載入,再好的選擇方法都繞不過這一點。
你最終會得到一個 suggest() 函式,最多返回一個技能名;一個
suggestion_block(),把它包起來放進系統提示詞;以及生成上面那張表的
harness,可以隨時指向你自己的名冊。
flowchart LR
subgraph C1["Call 1 - skim all 182 skills"]
direction TB
Q1["<b>Choice:</b> which skill fits?<br/><i>all 182, one line each</i>"]
N1["<b>Nouls:</b> need a skill at all?<br/>· act on their stuff?<br/>· follow written steps?<br/>· or just talk?"]
%% invisible link: without an edge these two share a rank, which in a TB
%% subgraph puts them side by side instead of stacked
Q1 ~~~ N1
end
subgraph C2["Call 2 - read those 3 properly"]
direction TB
Q2["<b>Choice:</b> which of the 3?<br/><i>with real detail now</i>"]
N2["<b>Nouls:</b> does each one<br/>really do it?"]
Q2 ~~~ N2
end
REQ["the request"] --> C1
C1 -->|"top 3"| C2
C1 -->|"nothing<br/>applies"| STOP["suggest<br/>nothing"]
C2 -->|"none fit"| STOP
C2 -->|"a winner"| OUT["suggest<br/>the winner"]
環境準備
- 安裝 TypeSafe 客戶端、Anthropic 客戶端,以及共享的 cookbook 輔助工具。
- 設定一個 TypeSafe API key,再為被測的 agent 設一個 Anthropic key。
pip install anthropic matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
export TYPESAFE_API_KEY=your-key-here
export ANTHROPIC_API_KEY=your-key-here
注意: 下面的程式碼塊是同一個指令碼,按順序排列。想跟著做,就把它們按給出的順序 放進一個檔案裡。
快取結果
JsonCache 會把每次呼叫的結果按輸入存起來,所以重新執行會重放
下面這些數字,而不是去調哪個 API。刪掉 json_cache.json 就能跑即時版。
釋出這次執行用的是 jev-1.12 和 claude-haiku-4-5-20251001,渲染於 2026-07-31。
import json
import os
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from time import perf_counter
import anthropic
import matplotlib
import matplotlib.pyplot as plt
from matplotlib.ticker import PercentFormatter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient
matplotlib.use("Agg") # headless render
TYPESAFE_MODEL = "jev-1.12"
AGENT_MODEL = (
"claude-haiku-4-5-20251001" # the agent under test, pinned so scores are stable
)
SHORTLIST = 3 # candidates carried from the first request into the second
EXCERPT_CHARS = (
700 # SKILL.md characters each candidate brings; the roster file stores 1600
)
GATE_THRESHOLD = (
0.30 # mean of the three request nouls, below which nothing is suggested
)
FITS_THRESHOLD = (
0.30 # a shortlist whose best "does this fit" noul is under this is dropped
)
WORKERS = 8 # small pool: enough to keep a live run to minutes, gentle on rate limits
assert EXCERPT_CHARS <= 1600, (
"the shipped roster file stores 1600 body characters per skill"
)
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
agent = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "cache-only"))
json_cache = JsonCache(Path("json_cache.json"))
第 1 步:載入名冊
hermes_roster.json 裝著
NousResearch/hermes-agent(MIT)在一個
固定提交上的 182 個技能。每條記錄包含技能的名字和分類、索引裡展示的描述、
完整描述,以及它 SKILL.md 的開頭。
下面的索引,以及提示詞裡它在上面那段指令,都是從 Hermes 原樣抄來的。
ROSTER = json.loads(Path("hermes_roster.json").read_text(encoding="utf-8"))
BY_NAME = {skill["name"]: skill for skill in ROSTER}
# Verbatim from hermes-agent agent/prompt_builder.py:build_skills_system_prompt.
PREAMBLE = (
"## Skills (mandatory)\n"
"Before replying, scan the skills below. If a skill matches or is even partially relevant "
"to your task, you MUST load it with skill_view(name) and follow its instructions. "
"Err on the side of loading — it is always better to have context you don't need "
"than to miss critical steps, pitfalls, or established workflows. "
"Skills contain specialized knowledge — API endpoints, tool-specific commands, "
"and proven workflows that outperform general-purpose approaches. Load the skill "
"even if you think you could handle the task with basic tools like web_search or terminal. "
"Skills also encode the user's preferred approach, conventions, and quality standards "
"for tasks like code review, planning, and testing — load them even for tasks you "
"already know how to do, because the skill defines how it should be done here.\n"
"Whenever the user asks you to configure, set up, install, enable, disable, modify, "
"or troubleshoot Hermes Agent itself — its CLI, config, models, providers, tools, "
"skills, voice, gateway, plugins, or any feature — load the `hermes-agent` skill "
"first. It has the actual commands (e.g. `hermes config set …`, `hermes tools`, "
"`hermes setup`) so you don't have to guess or invent workarounds.\n"
"If a skill has issues, fix it with skill_manage(action='patch').\n"
"After difficult/iterative tasks, offer to save as a skill. "
"If a skill you loaded was missing steps, had wrong commands, or needed "
"pitfalls you discovered, update it before finishing.\n"
"\n"
)
FOOTER = "\n\nOnly proceed without loading a skill if genuinely none are relevant to the task."
IDENTITY = (
"You are Hermes, a capable AI assistant with access to tools and a library "
"of skills. You help the user with coding, research, and everyday tasks.\n\n"
)
def render_index() -> str:
"""The body of <available_skills>: skills grouped by category, both sorted by name."""
by_category = defaultdict(list)
for skill in ROSTER:
by_category[skill["category"]].append(skill)
lines = []
for category in sorted(by_category):
lines.append(f" {category}:")
for skill in sorted(by_category[category], key=lambda s: s["name"]):
lines.append(f" - {skill['name']}: {skill['description']}")
return "\n".join(lines)
CATALOG_PROMPT = (
IDENTITY
+ PREAMBLE
+ "<available_skills>\n"
+ render_index()
+ "\n</available_skills>"
+ FOOTER
)
widths = [len(skill["description"]) for skill in ROSTER]
print(f"{len(ROSTER)} skills in {len({s['category'] for s in ROSTER})} categories")
print(f"roster prompt: {len(CATALOG_PROMPT):,} characters")
print(
f"index description: {sum(widths) / len(widths):.0f} characters on average, "
f"{max(widths)} at most"
)
print("\none category, as the agent reads it:")
index_lines = render_index().splitlines()
start = index_lines.index(" apple:")
end = next(
i
for i in range(start + 1, len(index_lines))
if not index_lines[i].startswith(" ")
)
print("\n".join(index_lines[start:end]))
182 skills in 33 categories
roster prompt: 16,089 characters
index description: 54 characters on average, 60 at most
one category, as the agent reads it:
apple:
- apple-notes: Manage Apple Notes via memo CLI: create, search, edit.
- apple-reminders: Apple Reminders via remindctl: add, list, complete.
- findmy: Track Apple devices/AirTags via FindMy.app on macOS.
- imessage: Send and receive iMessages/SMS via the imsg CLI on macOS.
第 2 步:先單獨給 agent 打分
requests.json 裡有 488 條單輪請求,其中 315 條正好由一個技能覆蓋,
另外 173 條沒有任何技能覆蓋。
有覆蓋的請求由 Claude Sonnet 5 依據每個技能自己的 SKILL.md 寫出,所以
標籤可信,而請求也比使用者實際發來的更容易。
這 173 條無覆蓋的請求都是用來懲罰亂猜的:85 條日常請求、42 條沒有技能能服務的技術問題(解釋一下什麼是 monad),以及 46 條索要 某個名冊裡沒有的技能——比如在一個只覆蓋 X、別的什麼都沒有的名冊上 把這條發到 Mastodon。
打分只看 agent 的第一次響應。兩個數字都是錯誤率,所以各自越低 越好:
- 錯誤載入:在有覆蓋的請求裡,第一次
skill_view呼叫 不是覆蓋技能的那部分佔比。整輪什麼也沒載入也算一次失誤。 - 多餘載入:在無覆蓋的請求裡,agent 呼叫了
skill_view的那部分佔比。
REQUESTS = json.loads(Path("requests.json").read_text(encoding="utf-8"))
POSITIVES = [p for p in REQUESTS if p["gold"]]
NEGATIVES = [p for p in REQUESTS if not p["gold"]]
print(
f"{len(REQUESTS)} requests: {len(POSITIVES)} covered by a skill "
f"({len({p['gold'] for p in POSITIVES})} distinct skills), {len(NEGATIVES)} covered by none"
)
print(f"\ncovered [{POSITIVES[0]['gold']}] {POSITIVES[0]['text']}")
print(f"uncovered {NEGATIVES[0]['text']}")
488 requests: 315 covered by a skill (171 distinct skills), 173 covered by none
covered [1password] I've got a config.yaml with `{{ op://app-prod/db/password }}` placeholders in it — can you set up my project to pull the real values in at runtime instead of hardcoding them?
uncovered Add these three cards to our Trello backlog.
建議放在系統提示詞單獨的一個塊裡,排在名冊之後而不是 名冊裡面,這樣名冊文本每一輪都完全一致,字首快取才保得住。
agent 有一套極簡的工具,包括用自由文本名載入技能的 skill_view。
名字必須和技能完全一致,才算一次正確載入。
# Verbatim from hermes-agent tools/skills_tool.py:SKILL_VIEW_SCHEMA.
SKILL_VIEW_DESCRIPTION = (
"Skills allow for loading information about specific tasks and workflows, as "
"well as scripts and templates. Load a skill's full content or access its "
"linked files (references, templates, scripts). First call returns SKILL.md "
"content plus a 'linked_files' dict showing available references/templates/"
"scripts. To access those, call again with file_path parameter."
)
TOOLS = [
{
"name": "skill_view",
"description": SKILL_VIEW_DESCRIPTION,
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string", "description": "The skill name."}
},
"required": ["name"],
},
},
{
"name": "terminal",
"description": "Run a shell command on the user's machine and return its output.",
"input_schema": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"],
},
},
{
"name": "read_file",
"description": "Read a file from the user's filesystem.",
"input_schema": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"],
},
},
{
"name": "web_search",
"description": "Search the web and return result snippets.",
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
]
@json_cache
def run_turn(model: str, arm: str, request: str, suggestion: str) -> dict:
"""One measured turn. ``arm`` is in the key so each arm samples independently."""
system = [
{"type": "text", "text": CATALOG_PROMPT, "cache_control": {"type": "ephemeral"}}
]
if suggestion:
system.append({"type": "text", "text": suggestion}) # after the breakpoint
response = agent.messages.create(
model=model,
max_tokens=1024,
system=system,
tools=TOOLS,
messages=[{"role": "user", "content": request}],
)
usage = response.usage
return {
"loaded": [
str(block.input.get("name", ""))
for block in response.content
if block.type == "tool_use" and block.name == "skill_view"
],
"input_tokens": usage.input_tokens or 0,
"output_tokens": usage.output_tokens or 0,
}
def summarise(turns: dict[str, dict]) -> dict[str, float]:
"""Two failure rates: wrong loads on covered requests, needless ones on uncovered."""
hits = [turns[p["text"]]["loaded"][:1] == [p["gold"]] for p in POSITIVES]
over = [bool(turns[p["text"]]["loaded"]) for p in NEGATIVES]
return {
# both metrics are errors, so the two columns read the same direction
"wrong_load": 1 - sum(hits) / len(hits),
"needless_load": sum(over) / len(over),
}
def run_arm(arm: str, suggestions: dict[str, str]) -> dict[str, dict]:
"""One measured turn per request, in a small pool. 488 calls."""
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
turns = pool.map(
lambda t: run_turn(AGENT_MODEL, arm, t, suggestions.get(t, "")), texts
)
return dict(zip(texts, turns))
agent 先只帶著自己的名冊跑一遍,也就是它今天的做法。它的兩個錯誤率 就是這份 cookbook 其餘部分拿來做對照的基線。
baseline = run_arm("baseline", {})
base_scores = summarise(baseline)
print(
f"wrong loads {base_scores['wrong_load']:.1%} ({len(POSITIVES)} covered requests)"
)
print(
f"needless loads {base_scores['needless_load']:.1%} ({len(NEGATIVES)} uncovered requests)"
)
# where the wrong loads land: a neighbour of the right skill, or somewhere unrelated?
misses = [
(p["gold"], baseline[p["text"]]["loaded"][0])
for p in POSITIVES
if baseline[p["text"]]["loaded"] and baseline[p["text"]]["loaded"][0] != p["gold"]
]
same_category = sum(
1
for gold, got in misses
if got in BY_NAME and BY_NAME[got]["category"] == BY_NAME[gold]["category"]
)
print(
f"\nof {len(misses)} wrong first picks, {same_category} came from the right skill's own "
f"category"
)
wrong loads 16.8% (315 covered requests)
needless loads 9.8% (173 uncovered requests)
of 36 wrong first picks, 10 came from the right skill's own category
錯誤載入落在正確技能自己分類裡的比例,遠高於隨機水平,所以難的地方 在於把幾個長得像的區分開。agent 已經在差不多對的地方找了。
第 3 步:給整個名冊排序
一次請求帶著兩類問題:
which是一道Choice問題, 覆蓋全部 182 個技能名,以索引描述作為每個選項的 criteria(也就是 agent 自己拿到的那段文本)。它的機率就是排序結果。- 三道關於請求的
Noul問題, 列在下面,各自用不同的方式問:它想要的是執行動作,還是給出解釋。prose_suffices的方向相反。它們的均值決定要不要給出任何建議, 低於 0.30 就什麼都不建議。
兩者在同一次請求裡發出,所以排序和檢查只花一個來回。
這三道題要寫成在問「是否想要一個動作」。一個關於主題內容的問題 是沒法把解釋一下什麼是 monad 和一個真需要技能的請求區分開的,因為兩者 都跟軟體有關。
一道 Choice 問題裝下這麼大的名冊很輕鬆。再大幾倍,你就該
把它拆成幾塊分別排序,再對這一批勝出者跑同樣的短名單步驟。
CHOICE_INSTRUCTIONS = (
"Which of these skills, if any, is the right one to load to help with the "
"user's latest request?"
)
GATE_QUESTIONS = {
"acts_on_user_system": (
"Is the assistant being asked to act on the user's files, accounts, devices, "
"or online services, rather than only to explain or advise?"
),
"would_follow_documented_procedure": (
"Would a careful expert answering this consult a specific documented procedure "
"or set of commands, rather than answering from general understanding?"
),
"prose_suffices": (
"Could a knowledgeable generalist fully satisfy this request in prose, with "
"no tools, no documentation, and no access to the user's files or accounts?"
),
}
INVERTED = {"prose_suffices"} # a yes here points away from needing a skill
def build_state(request: str) -> dict:
return {"request": request, "recent_context": ""}
@json_cache
def rank_wide(request: str) -> dict:
"""Request 1: rank all 182 skills, and score the request for whether a skill applies."""
questions = {
"which": Choice(
instructions=CHOICE_INSTRUCTIONS,
criteria={skill["name"]: skill["description"] for skill in ROSTER},
)
}
for key, text in GATE_QUESTIONS.items():
questions[f"gate::{key}"] = Noul(instructions=text)
started = perf_counter()
response = client.system_one(
state=build_state(request), questions=questions, model=TYPESAFE_MODEL
)
ranked = sorted(
response.answers["which"].probabilities.items(), key=lambda kv: -kv[1]
)
values = {
key.removeprefix("gate::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("gate::")
}
oriented = [(1.0 - v) if k in INVERTED else v for k, v in values.items()]
return {
"ranked": ranked[
:12
], # more than any shortlist needs, and keeps the cache small
"gate": sum(oriented) / len(oriented),
"values": values,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
DEMO = [
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so it syncs"
" to my phone? Just write it up in whatever editor pops up.",
"Can you put together a pitch deck skeleton (cover, situation overview, comps, precedent"
" transactions, DCF, LBO) as a .pptx, using our firm-template.pptx for branding and"
" footnoting each valuation number back to the cell it came from in the model?",
"Post this announcement to my Mastodon account.",
]
for request in DEMO:
wide = rank_wide(request)
verdict = "suggest" if wide["gate"] >= GATE_THRESHOLD else "stay quiet"
print(f'"{request[:78]}"')
print(f" needs a skill {wide['gate']:.2f} -> {verdict} ({wide['seconds']}s)")
for name, probability in wide["ranked"][:SHORTLIST]:
print(f" {probability:.3f} {name:<38}{BY_NAME[name]['description']}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
needs a skill 0.75 -> suggest (0.31s)
0.990 apple-notes Manage Apple Notes via memo CLI: create, search, edit.
0.010 computer-use Drive the user's desktop in the background — clicking, ty...
0.000 concept-diagrams Generate flat, minimal educational SVG visuals as HTML.
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
needs a skill 0.76 -> suggest (0.16s)
0.700 powerpoint Create, read, edit .pptx decks, slides, notes, templates.
0.300 pptx-author Build PowerPoint decks headless with python-pptx.
0.000 chroma Embedding database for RAG and semantic search.
"Post this announcement to my Mastodon account."
needs a skill 0.78 -> suggest (0.16s)
0.550 xurl X/Twitter via xurl CLI: raw post search, posting, DM, media.
0.140 computer-use Drive the user's desktop in the background — clicking, ty...
0.080 openhands Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).
Notes.app 那條請求沒有歧義,排在第一的選項就是對的。Mastodon 那條 則無藥可救:三道題都說需要一個技能,因為往一個賬號發帖是一個 動作;而名冊裡有發到 X 的技能、沒有發到 Mastodon 的,最接近的那個技能 照樣會勝出。
剩下那份 deck。領先的兩個都是 .pptx 技能,而在 60 個字元下,這道寬範圍 Choice
問題把編輯技能排在創作技能前面——儘管請求要的是
創作一份 deck。
第 4 步:重排前三名
三個選項留出了空間,可以放下完整描述再加上每個技能自己
SKILL.md 的開頭,於是第二次請求把同一個問題放到更好的證據上:
which是一道覆蓋短名單的Choice問題,以那段更長的文本作為每個 選項的 criteria。fits::{name}是每個候選一道Noul問題:這個技能是不是正好 做請求要的那件事?每道都獨立作答,所以它們可能全都偏低; 如果短名單裡最高的那個低於 0.30,整份短名單就被丟掉。
RERANK_INSTRUCTIONS = (
"Exactly one of these skills is the right one to load for the user's latest "
"request. Which one? Read what each actually does, not just its name."
)
def rerank_criteria(names: tuple[str, ...], excerpt: int) -> dict[str, str]:
return {
name: f"{BY_NAME[name]['description_full']} — {BY_NAME[name]['body'][:excerpt]}"
for name in names
}
def rerank_questions(names: tuple[str, ...], excerpt: int) -> dict:
questions = {
"which": Choice(
instructions=RERANK_INSTRUCTIONS, criteria=rerank_criteria(names, excerpt)
)
}
for name in names:
questions[f"fits::{name}"] = Noul(
instructions=(
f"Does the skill '{name}' do the specific thing the user's request asks "
f"for? It is described as: {BY_NAME[name]['description_full']}"
)
)
return questions
@json_cache
def rerank(request: str, names: tuple[str, ...], excerpt: int) -> dict:
"""Request 2: the same Choice over a shortlist, plus one absolute noul per candidate."""
started = perf_counter()
response = client.system_one(
state=build_state(request),
questions=rerank_questions(names, excerpt),
model=TYPESAFE_MODEL,
)
return {
"winner": response.answers["which"].choice,
"fits": {
key.removeprefix("fits::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("fits::")
},
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
for request in DEMO:
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
print(f'"{request[:78]}"\n scored too low, nothing suggested\n')
continue
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
best = max(result["fits"].values())
verdict = result["winner"] if best >= FITS_THRESHOLD else "nothing fits"
print(f'"{request[:78]}"')
print(f" was {shortlist[0]} -> {verdict} ({result['seconds']}s)")
for name in shortlist:
print(f" fits {result['fits'][name]:.2f} {name}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
was apple-notes -> apple-notes (0.12s)
fits 0.60 apple-notes
fits 0.54 computer-use
fits 0.01 concept-diagrams
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
was powerpoint -> pptx-author (0.09s)
fits 0.73 powerpoint
fits 0.38 pptx-author
fits 0.02 chroma
"Post this announcement to my Mastodon account."
was xurl -> xurl (0.09s)
fits 0.56 xurl
fits 0.38 computer-use
fits 0.05 openhands
兩個 .pptx 技能在各自帶上自己的文本後就分開了:那份 deck 請求改判
給創作技能。
這裡的 fits noul 和 Choice 是分歧的:noul 給編輯技能打的分更高,
而 Choice 選的是創作技能。它們在決定的是不同的事。Choice
定的是哪一個技能,noul 定的是到底要不要說點什麼。
Mastodon 那條請求兩道檢查都通過了:它最好的 fits noul 高於 0.30,於是這份
方案為一條關於 Mastodon 的請求建議了 X 技能。大多數類似的請求都能被抓出來。
第二輪只能拒絕寬範圍排序交給它的東西,而這裡交給它的是三個
擦邊球。
下面這個函式就是整套方案:兩次請求、兩個閾值,最多返回一個 技能名。
要把它指向你自己的名冊,替換 hermes_roster.json 即可。上面每個問題都只
從這個檔案裡讀 name、description、description_full 和 body,除此之外沒有別的地方
知道 Hermes。
def suggest(request: str) -> tuple[str, ...]:
"""At most one skill name for a request, or () for "nothing here applies"."""
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
return ()
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
if max(result["fits"].values()) < FITS_THRESHOLD:
return ()
return (result["winner"],)
def suggestion_block(names: tuple[str, ...]) -> str:
"""What gets appended after the roster, in the suggestion.
This string is a measured input rather than prose: it goes to the agent, so it is part
of every graded turn's cache key. Editing a word here silently invalidates the shipped
results and costs a live re-run to restore them.
"""
body = (
f"Relevant to the current request: {', '.join(names)}. Ignore this if it does not "
"fit what the user actually asked for."
if names
else "No skill in the roster appears relevant to this request."
)
return f"\n\n<skill_relevance>\n{body}\n</skill_relevance>"
print(suggestion_block(suggest(DEMO[1])))
print(suggestion_block(suggest(DEMO[2])))
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
<skill_relevance>
Relevant to the current request: xurl. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
第 5 步:衡量建議的效果
488 條請求各自送給 agent 三次,每次一個被測輪次。各次執行 只差在告訴 agent 什麼:
| 系統提示詞裡放什麼 | |
|---|---|
| 單獨的 agent | 什麼都不放 |
| 帶建議的 agent | suggest() 返回的東西 |
| 直接拿到答案的 agent | 覆蓋技能的名字;沒有覆蓋技能時則是「沒有適用的」 |
第三行是做不到的;它是另外兩行拿來對照的天花板。
那句建議的措辭同時在幹兩件事。它說明建議可以忽略, 因為推得越用力,錯誤建議也越會被照辦,而一條錯誤建議比沒有建議更糟。而且, 一輪裡即使沒有任何可建議的東西,也仍然會發一句說明;什麼都不發 的話,名冊自己那句「寧可多載入」的指令就沒人制衡了。
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool: # up to 488 x 2 TypeSafe requests
suggested = dict(zip(texts, pool.map(suggest, texts)))
WIDE = {text: rank_wide(text) for text in texts} # all cache hits now; reused below
arms = {
"baseline": {},
"TypeSafe": {
request["text"]: suggestion_block(suggested[request["text"]])
for request in REQUESTS
},
"oracle": {
request["text"]: suggestion_block((request["gold"],) if request["gold"] else ())
for request in REQUESTS
},
}
scores = {
arm: summarise(run_arm(arm, suggestions)) for arm, suggestions in arms.items()
}
print(f"{'run':<10}{'wrong loads':>13}{'needless loads':>16}")
for arm, row in scores.items():
print(f"{arm:<10}{row['wrong_load']:>13.1%}{row['needless_load']:>16.1%}")
def fewer(metric: str) -> str:
"""The plain ratio between the two arms' error rates."""
return f"{scores['baseline'][metric] / scores['TypeSafe'][metric]:.1f}x fewer"
print(
f"\nbaseline -> TypeSafe: {fewer('wrong_load')} wrong loads, "
f"{fewer('needless_load')} needless ones"
)
run wrong loads needless loads
baseline 16.8% 9.8%
TypeSafe 7.3% 4.0%
oracle 2.5% 1.2%
baseline -> TypeSafe: 2.3x fewer wrong loads, 2.4x fewer needless ones
moved = [
(
baseline[p["text"]]["loaded"][:1] == [p["gold"]],
run_turn(AGENT_MODEL, "TypeSafe", p["text"], arms["TypeSafe"][p["text"]])[
"loaded"
][:1]
== [p["gold"]],
)
for p in POSITIVES
]
print(
f"of {len(POSITIVES)} covered requests: {sum(not b and a for b, a in moved)} the suggestion "
f"fixed, {sum(b and not a for b, a in moved)} it broke"
)
of 315 covered requests: 37 the suggestion fixed, 7 it broke
建議修好的請求遠多於它弄壞的,但它確實也弄壞了一些 agent 本來 答對的。一條自信的錯誤建議比完全沒有建議更有說服力, 這就是把建議放到這一輪前面的代價。
SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"
ARM_COLOR = {"baseline": BLUE, "TypeSafe": ORANGE, "oracle": MUTED}
def style(ax):
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
panels = [
("wrong_load", f"wrong loads\n{len(POSITIVES)} covered requests"),
("needless_load", f"needless loads\n{len(NEGATIVES)} uncovered requests"),
]
names = list(scores)
fig, axes = plt.subplots(1, 2, figsize=(8.4, 3.6), facecolor=SURFACE)
for ax, (metric, title) in zip(axes, panels):
style(ax)
ax.grid(axis="y", color=GRID, linewidth=0.8)
values = [scores[arm][metric] for arm in names]
bars = ax.bar(
names,
values,
0.58,
color=[ARM_COLOR[arm] for arm in names],
# the oracle is a ceiling, not a competitor: gray, and hatched so it never depends
# on colour alone
hatch=["", "", "///"],
edgecolor=SURFACE,
linewidth=1.2,
)
ax.bar_label(
bars,
labels=[f"{v:.1%}" for v in values],
padding=3,
color=INK2,
fontsize=9,
)
ax.set_title(title, loc="left", color=INK2, fontsize=9.5)
ax.set_ylim(0, max(values) * 1.28)
ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))
ax.set_ylabel("% of those requests - lower is better", color=INK2, fontsize=9)
fig.suptitle(
f"Hermes' {len(ROSTER)}-skill roster, {len(REQUESTS)} requests, {AGENT_MODEL}",
x=0.02,
ha="left",
color=INK,
fontsize=11,
)
fig.tight_layout()
display(fig)
plt.close(fig)
結果說明了什麼
- 錯誤載入從 16.8% 降到 7.3%,多餘載入從 9.8% 降到 4.0%,這差不多填上了 「從截斷的索引去猜」和「直接拿到答案」之間的大部分差距。
- 有些 agent 本來自己答對的請求,加上建議後反而錯了。具體數量見上。
當你的 agent 帶著一份龐大的名冊時,就照這個形狀來:先對所有技能做一次 廉價的排序,再仔細看兩三個。每一步都有可能空手而歸。
在 playground 裡開啟它
為第 4 步裡那份 deck 請求生成一個 playground 連結,用每個候選的完整 描述和正文節選作為它的 criteria。
demo_shortlist = tuple(name for name, _ in rank_wide(DEMO[1])["ranked"][:SHORTLIST])
playground_link = make_playground_link(
build_state(DEMO[1]),
rerank_questions(demo_shortlist, EXCERPT_CHARS),
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open the shortlist + questions in the TypeSafe playground]({playground_link})"
)
)
在 TypeSafe playground 中開啟短名單 + 問題 →
接下來
同樣的形狀在別處也會出現: 意圖路由是用來路由到 處理器而不是技能的,置信度講 怎麼選那兩個閾值, 推測式扇出講怎麼把每個 問題放進一次請求。