スキルの提案
スキルの提案
Hermes のカタログにある 182 個のスキルから、1 ターンにつき最大 1 つを選びます。2 回の TypeSafe リクエストで上位候補を順位付けし、再確認します。
エージェントはスキルを切り詰めてすべてシステムメッセージに読み込むことで選択しますが、 これはコストを増やし、スキル選択の性能を劣化させ、セッションの残りにコンテキストの 腐敗を引き起こします。これを、1 ターンにつき 2 回の TypeSafe リクエスト(一方はスキルの 順位付け、もう一方は選択の検証)で解決し、誤ったスキルの読み込みを半分以下に減らします。
大きなスキル名簿を持つエージェントは、ほとんど情報がないまま選択を行います。名簿は
インデックスとして届きます。スキルごとに 1 行で、説明は全文が会話を圧迫しないよう
切り詰められています。ここで使うエージェント harness である Hermes は、既定で 60 文字に
切り詰めます。たとえばその幅では、.pptx ファイルを 編集する スキルと 作成する
スキルはほぼ同じに読めます。ピッチデックを頼むと、間違った方を読み込むかもしれません。
どのスキルも当てはまらないターンでも、それでも 1 つ読み込んでしまうことがあります。
名前のリストは当て推量を誘うからです。
この cookbook は説明に手を付けず、代わりに段階的な開示を使います。182 個のスキルを 安価に読み、その後で 3 つを詳しく読みます。どのスキルを読み込むか(あるいは読み込まない か)の判断の前に、2 回の TypeSafe リクエストを置きます。1 回目は名簿のすべてのスキルを ユーザーのターンに対して順位付けし、そのターンにスキルが必要かどうかにも答えます。 2 回目は上位 3 つだけを読み直します。今度は各スキルの完全な説明と instructions の 冒頭を伴い、それらすべてを却下する自由もあります。
勝者の名前は、そのターンのエージェントのシステムプロンプトに 1 行だけ追加されます。
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user
actually asked for.
</skill_relevance>
エージェントは完全なインデックスと自身の判断を保ち、その 1 行はどのエントリを最初に
見るかだけを伝えます。名簿自体は決して変わらないので、それに対するプレフィックス
キャッシュはそのまま有効です。claude-haiku-4-5-20251001 に対する 488 件のリクエストで、
Hermes の名簿のスキルを使った結果:
| 誤ったスキルを読み込む | 何も合わないのに読み込む | |
|---|---|---|
| 名簿だけを持つエージェント単体 | 16.8% | 9.8% |
| TypeSafe の提案を受けたエージェント | 7.3% | 4.0% |
| 正解を渡されたエージェント | 2.5% | 1.2% |
3 行目は、間違いの下限がゼロではないことを示します。正しいスキルを渡されたエージェントでも 常にそれを読み込むわけではなく、どんなに優れた選択方法でもそれを超えられません。
最終的には、最大 1 つのスキル名を返す suggest() 関数、それをシステムプロンプト用に
包む suggestion_block()、そして上の表を生み出した harness が得られ、自分の名簿に
向けられるようになります。
flowchart LR
subgraph C1["Call 1 - skim all 182 skills"]
direction TB
Q1["<b>Choice:</b> which skill fits?<br/><i>all 182, one line each</i>"]
N1["<b>Nouls:</b> need a skill at all?<br/>· act on their stuff?<br/>· follow written steps?<br/>· or just talk?"]
%% invisible link: without an edge these two share a rank, which in a TB
%% subgraph puts them side by side instead of stacked
Q1 ~~~ N1
end
subgraph C2["Call 2 - read those 3 properly"]
direction TB
Q2["<b>Choice:</b> which of the 3?<br/><i>with real detail now</i>"]
N2["<b>Nouls:</b> does each one<br/>really do it?"]
Q2 ~~~ N2
end
REQ["the request"] --> C1
C1 -->|"top 3"| C2
C1 -->|"nothing<br/>applies"| STOP["suggest<br/>nothing"]
C2 -->|"none fit"| STOP
C2 -->|"a winner"| OUT["suggest<br/>the winner"]
準備
- TypeSafe クライアント、Anthropic クライアント、共有の cookbook ヘルパーを インストールします。
- TypeSafe API キー と、測定対象のエージェント用の Anthropic キーを設定します。
pip install anthropic matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
export TYPESAFE_API_KEY=your-key-here
export ANTHROPIC_API_KEY=your-key-here
Note: 以下のコードブロックは順番どおりの 1 つのスクリプトです。追いかけるには、 示された順番で 1 つのファイルに入れます。
結果のキャッシュ
JsonCache は各呼び出しの結果をその入力に基づいて保存するので、再実行するとどちらの
API も呼ばずに以下の数値を再生します。実際に実行するには json_cache.json を削除します。
公開された実行は jev-1.12 と claude-haiku-4-5-20251001 を使い、2026-07-31 に描画
されました。
import json
import os
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from time import perf_counter
import anthropic
import matplotlib
import matplotlib.pyplot as plt
from matplotlib.ticker import PercentFormatter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient
matplotlib.use("Agg") # headless render
TYPESAFE_MODEL = "jev-1.12"
AGENT_MODEL = (
"claude-haiku-4-5-20251001" # the agent under test, pinned so scores are stable
)
SHORTLIST = 3 # candidates carried from the first request into the second
EXCERPT_CHARS = (
700 # SKILL.md characters each candidate brings; the roster file stores 1600
)
GATE_THRESHOLD = (
0.30 # mean of the three request nouls, below which nothing is suggested
)
FITS_THRESHOLD = (
0.30 # a shortlist whose best "does this fit" noul is under this is dropped
)
WORKERS = 8 # small pool: enough to keep a live run to minutes, gentle on rate limits
assert EXCERPT_CHARS <= 1600, (
"the shipped roster file stores 1600 body characters per skill"
)
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
agent = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "cache-only"))
json_cache = JsonCache(Path("json_cache.json"))
ステップ 1:名簿を読み込む
hermes_roster.json には、ある固定したコミット時点の
NousResearch/hermes-agent(MIT)の
182 個のスキルが入っています。各レコードは、スキルの名前とカテゴリ、インデックスが
表示する説明、完全な説明、そして SKILL.md の冒頭を保持します。
下のインデックスと、プロンプト内でその上にある instructions は、Hermes からそのまま コピーしたものです。
ROSTER = json.loads(Path("hermes_roster.json").read_text(encoding="utf-8"))
BY_NAME = {skill["name"]: skill for skill in ROSTER}
# Verbatim from hermes-agent agent/prompt_builder.py:build_skills_system_prompt.
PREAMBLE = (
"## Skills (mandatory)\n"
"Before replying, scan the skills below. If a skill matches or is even partially relevant "
"to your task, you MUST load it with skill_view(name) and follow its instructions. "
"Err on the side of loading — it is always better to have context you don't need "
"than to miss critical steps, pitfalls, or established workflows. "
"Skills contain specialized knowledge — API endpoints, tool-specific commands, "
"and proven workflows that outperform general-purpose approaches. Load the skill "
"even if you think you could handle the task with basic tools like web_search or terminal. "
"Skills also encode the user's preferred approach, conventions, and quality standards "
"for tasks like code review, planning, and testing — load them even for tasks you "
"already know how to do, because the skill defines how it should be done here.\n"
"Whenever the user asks you to configure, set up, install, enable, disable, modify, "
"or troubleshoot Hermes Agent itself — its CLI, config, models, providers, tools, "
"skills, voice, gateway, plugins, or any feature — load the `hermes-agent` skill "
"first. It has the actual commands (e.g. `hermes config set …`, `hermes tools`, "
"`hermes setup`) so you don't have to guess or invent workarounds.\n"
"If a skill has issues, fix it with skill_manage(action='patch').\n"
"After difficult/iterative tasks, offer to save as a skill. "
"If a skill you loaded was missing steps, had wrong commands, or needed "
"pitfalls you discovered, update it before finishing.\n"
"\n"
)
FOOTER = "\n\nOnly proceed without loading a skill if genuinely none are relevant to the task."
IDENTITY = (
"You are Hermes, a capable AI assistant with access to tools and a library "
"of skills. You help the user with coding, research, and everyday tasks.\n\n"
)
def render_index() -> str:
"""The body of <available_skills>: skills grouped by category, both sorted by name."""
by_category = defaultdict(list)
for skill in ROSTER:
by_category[skill["category"]].append(skill)
lines = []
for category in sorted(by_category):
lines.append(f" {category}:")
for skill in sorted(by_category[category], key=lambda s: s["name"]):
lines.append(f" - {skill['name']}: {skill['description']}")
return "\n".join(lines)
CATALOG_PROMPT = (
IDENTITY
+ PREAMBLE
+ "<available_skills>\n"
+ render_index()
+ "\n</available_skills>"
+ FOOTER
)
widths = [len(skill["description"]) for skill in ROSTER]
print(f"{len(ROSTER)} skills in {len({s['category'] for s in ROSTER})} categories")
print(f"roster prompt: {len(CATALOG_PROMPT):,} characters")
print(
f"index description: {sum(widths) / len(widths):.0f} characters on average, "
f"{max(widths)} at most"
)
print("\none category, as the agent reads it:")
index_lines = render_index().splitlines()
start = index_lines.index(" apple:")
end = next(
i
for i in range(start + 1, len(index_lines))
if not index_lines[i].startswith(" ")
)
print("\n".join(index_lines[start:end]))
182 skills in 33 categories
roster prompt: 16,089 characters
index description: 54 characters on average, 60 at most
one category, as the agent reads it:
apple:
- apple-notes: Manage Apple Notes via memo CLI: create, search, edit.
- apple-reminders: Apple Reminders via remindctl: add, list, complete.
- findmy: Track Apple devices/AirTags via FindMy.app on macOS.
- imessage: Send and receive iMessages/SMS via the imsg CLI on macOS.
ステップ 2:エージェント単体を採点する
requests.json には 488 件の単一ターンのリクエストが入っており、そのうち 315 件は
ちょうど 1 つのスキルにカバーされ、残りの 173 件は何にもカバーされていません。
カバーされたリクエストは、各スキル自身の SKILL.md から Claude Sonnet 5 が書いたもの
なので、ラベルは信頼でき、リクエストはユーザーが送るものより簡単です。
カバーされていない 173 件はすべて、当て推量に罰を与えるために書かれました。日常的な リクエストが 85 件、どのスキルも対応しない技術的な質問が 42 件(モナドとは何かを 説明して)、そして名簿にスキルのない具体的な何かを求めるものが 46 件です。たとえば、 X だけをカバーして他に何もない名簿に これを Mastodon に投稿して といった具合です。
採点はエージェントの最初の応答だけを読み取ります。どちらの数値もエラー率なので、 低いほど良いです。
- wrong load:カバーされたリクエストのうち、最初の
skill_view呼び出しが正しい スキルではなかった割合。何も読み込まなかったターンもミスとして数えます。 - needless load:カバーされていないリクエストのうち、エージェントが
skill_viewを 呼んだ割合。
REQUESTS = json.loads(Path("requests.json").read_text(encoding="utf-8"))
POSITIVES = [p for p in REQUESTS if p["gold"]]
NEGATIVES = [p for p in REQUESTS if not p["gold"]]
print(
f"{len(REQUESTS)} requests: {len(POSITIVES)} covered by a skill "
f"({len({p['gold'] for p in POSITIVES})} distinct skills), {len(NEGATIVES)} covered by none"
)
print(f"\ncovered [{POSITIVES[0]['gold']}] {POSITIVES[0]['text']}")
print(f"uncovered {NEGATIVES[0]['text']}")
488 requests: 315 covered by a skill (171 distinct skills), 173 covered by none
covered [1password] I've got a config.yaml with `{{ op://app-prod/db/password }}` placeholders in it — can you set up my project to pull the real values in at runtime instead of hardcoding them?
uncovered Add these three cards to our Trello backlog.
提案はシステムプロンプトの独自のブロックに入り、名簿の中ではなく後に置かれます。そう すればプレフィックスキャッシュを保つために、名簿のテキストが毎ターン同一になります。
エージェントは最小限のツールセットを持ち、その中には自由テキストの名前でスキルを
読み込む skill_view が含まれます。正しく読み込むには、名前がスキルと完全に一致する
必要があります。
# Verbatim from hermes-agent tools/skills_tool.py:SKILL_VIEW_SCHEMA.
SKILL_VIEW_DESCRIPTION = (
"Skills allow for loading information about specific tasks and workflows, as "
"well as scripts and templates. Load a skill's full content or access its "
"linked files (references, templates, scripts). First call returns SKILL.md "
"content plus a 'linked_files' dict showing available references/templates/"
"scripts. To access those, call again with file_path parameter."
)
TOOLS = [
{
"name": "skill_view",
"description": SKILL_VIEW_DESCRIPTION,
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string", "description": "The skill name."}
},
"required": ["name"],
},
},
{
"name": "terminal",
"description": "Run a shell command on the user's machine and return its output.",
"input_schema": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"],
},
},
{
"name": "read_file",
"description": "Read a file from the user's filesystem.",
"input_schema": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"],
},
},
{
"name": "web_search",
"description": "Search the web and return result snippets.",
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
]
@json_cache
def run_turn(model: str, arm: str, request: str, suggestion: str) -> dict:
"""One measured turn. ``arm`` is in the key so each arm samples independently."""
system = [
{"type": "text", "text": CATALOG_PROMPT, "cache_control": {"type": "ephemeral"}}
]
if suggestion:
system.append({"type": "text", "text": suggestion}) # after the breakpoint
response = agent.messages.create(
model=model,
max_tokens=1024,
system=system,
tools=TOOLS,
messages=[{"role": "user", "content": request}],
)
usage = response.usage
return {
"loaded": [
str(block.input.get("name", ""))
for block in response.content
if block.type == "tool_use" and block.name == "skill_view"
],
"input_tokens": usage.input_tokens or 0,
"output_tokens": usage.output_tokens or 0,
}
def summarise(turns: dict[str, dict]) -> dict[str, float]:
"""Two failure rates: wrong loads on covered requests, needless ones on uncovered."""
hits = [turns[p["text"]]["loaded"][:1] == [p["gold"]] for p in POSITIVES]
over = [bool(turns[p["text"]]["loaded"]) for p in NEGATIVES]
return {
# both metrics are errors, so the two columns read the same direction
"wrong_load": 1 - sum(hits) / len(hits),
"needless_load": sum(over) / len(over),
}
def run_arm(arm: str, suggestions: dict[str, str]) -> dict[str, dict]:
"""One measured turn per request, in a small pool. 488 calls."""
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
turns = pool.map(
lambda t: run_turn(AGENT_MODEL, arm, t, suggestions.get(t, "")), texts
)
return dict(zip(texts, turns))
まずエージェントは名簿だけを持って実行します。現在の動作どおりです。その 2 つのエラー率が、 cookbook の残りが比較するベースラインになります。
baseline = run_arm("baseline", {})
base_scores = summarise(baseline)
print(
f"wrong loads {base_scores['wrong_load']:.1%} ({len(POSITIVES)} covered requests)"
)
print(
f"needless loads {base_scores['needless_load']:.1%} ({len(NEGATIVES)} uncovered requests)"
)
# where the wrong loads land: a neighbour of the right skill, or somewhere unrelated?
misses = [
(p["gold"], baseline[p["text"]]["loaded"][0])
for p in POSITIVES
if baseline[p["text"]]["loaded"] and baseline[p["text"]]["loaded"][0] != p["gold"]
]
same_category = sum(
1
for gold, got in misses
if got in BY_NAME and BY_NAME[got]["category"] == BY_NAME[gold]["category"]
)
print(
f"\nof {len(misses)} wrong first picks, {same_category} came from the right skill's own "
f"category"
)
wrong loads 16.8% (315 covered requests)
needless loads 9.8% (173 uncovered requests)
of 36 wrong first picks, 10 came from the right skill's own category
誤った読み込みは、偶然よりはるかに高い頻度で正しいスキルと同じカテゴリに落ちます。 そのため難しいのは、いくつかの似た者同士を見分けることです。エージェントはすでに おおよそ正しい場所を見ています。
ステップ 3:名簿全体を順位付けする
1 回のリクエストが 2 種類の質問を運びます。
whichは 182 個すべてのスキル名に対するChoiceの質問で、 インデックスの説明が各選択肢の criteria になります(エージェント自身が受け取るのと 同じテキストです)。その確率が順位になります。- リクエストについての 3 つの
Noulの質問で、以下に示します。 それぞれが、説明を与えるのではなく行動を求めるかどうかを別の言い方で尋ねます。prose_sufficesは逆に数えます。その平均が、そもそも何かを提案するかどうかを決め、 0.30 未満なら何も提案しません。
どちらも 1 回のリクエストで送られるので、順位付けとチェックは往復 1 回で済みます。
この 3 つは、行動が求められているかを尋ねるように書きます。主題についての質問では、 モナドとは何かを説明して とスキルを必要とするリクエストを区別できません。どちらも ソフトウェアだからです。
Choice の質問 1 つで、この規模の名簿は余裕を持って収まります。数倍大きくなったら、
チャンクに分割してそれぞれを順位付けし、同じ候補絞り込みのステップを勝者に対して
実行します。
CHOICE_INSTRUCTIONS = (
"Which of these skills, if any, is the right one to load to help with the "
"user's latest request?"
)
GATE_QUESTIONS = {
"acts_on_user_system": (
"Is the assistant being asked to act on the user's files, accounts, devices, "
"or online services, rather than only to explain or advise?"
),
"would_follow_documented_procedure": (
"Would a careful expert answering this consult a specific documented procedure "
"or set of commands, rather than answering from general understanding?"
),
"prose_suffices": (
"Could a knowledgeable generalist fully satisfy this request in prose, with "
"no tools, no documentation, and no access to the user's files or accounts?"
),
}
INVERTED = {"prose_suffices"} # a yes here points away from needing a skill
def build_state(request: str) -> dict:
return {"request": request, "recent_context": ""}
@json_cache
def rank_wide(request: str) -> dict:
"""Request 1: rank all 182 skills, and score the request for whether a skill applies."""
questions = {
"which": Choice(
instructions=CHOICE_INSTRUCTIONS,
criteria={skill["name"]: skill["description"] for skill in ROSTER},
)
}
for key, text in GATE_QUESTIONS.items():
questions[f"gate::{key}"] = Noul(instructions=text)
started = perf_counter()
response = client.system_one(
state=build_state(request), questions=questions, model=TYPESAFE_MODEL
)
ranked = sorted(
response.answers["which"].probabilities.items(), key=lambda kv: -kv[1]
)
values = {
key.removeprefix("gate::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("gate::")
}
oriented = [(1.0 - v) if k in INVERTED else v for k, v in values.items()]
return {
"ranked": ranked[
:12
], # more than any shortlist needs, and keeps the cache small
"gate": sum(oriented) / len(oriented),
"values": values,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
DEMO = [
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so it syncs"
" to my phone? Just write it up in whatever editor pops up.",
"Can you put together a pitch deck skeleton (cover, situation overview, comps, precedent"
" transactions, DCF, LBO) as a .pptx, using our firm-template.pptx for branding and"
" footnoting each valuation number back to the cell it came from in the model?",
"Post this announcement to my Mastodon account.",
]
for request in DEMO:
wide = rank_wide(request)
verdict = "suggest" if wide["gate"] >= GATE_THRESHOLD else "stay quiet"
print(f'"{request[:78]}"')
print(f" needs a skill {wide['gate']:.2f} -> {verdict} ({wide['seconds']}s)")
for name, probability in wide["ranked"][:SHORTLIST]:
print(f" {probability:.3f} {name:<38}{BY_NAME[name]['description']}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
needs a skill 0.75 -> suggest (0.31s)
0.990 apple-notes Manage Apple Notes via memo CLI: create, search, edit.
0.010 computer-use Drive the user's desktop in the background — clicking, ty...
0.000 concept-diagrams Generate flat, minimal educational SVG visuals as HTML.
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
needs a skill 0.76 -> suggest (0.16s)
0.700 powerpoint Create, read, edit .pptx decks, slides, notes, templates.
0.300 pptx-author Build PowerPoint decks headless with python-pptx.
0.000 chroma Embedding database for RAG and semantic search.
"Post this announcement to my Mastodon account."
needs a skill 0.78 -> suggest (0.16s)
0.550 xurl X/Twitter via xurl CLI: raw post search, posting, DM, media.
0.140 computer-use Drive the user's desktop in the background — clicking, ty...
0.080 openhands Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).
Notes.app のリクエストは明確で、その最上位の選択肢が正解です。Mastodon のものは 順位付けでは救えません。3 つの質問はスキルが求められていると答え、アカウントへの投稿は 行動だからです。X への投稿用のスキルはあっても Mastodon 用は何もないので、どのみち 最も近いスキルが勝ちます。
残るはデックです。上位 2 つはどちらも .pptx スキルで、60 文字では、広い Choice の
質問は、デックの作成についてのリクエストに対して、作成スキルより編集スキルを上位に
置きます。
ステップ 4:上位 3 つを再ランキングする
3 つの選択肢なら、完全な説明に加えて各スキル自身の SKILL.md の冒頭を入れる余地が
あるので、2 回目のリクエストは同じ質問をより良い証拠に投げかけます。
whichは候補リストに対するChoiceの質問で、その長いテキストが各選択肢の criteria になります。fits::{name}は候補ごとのNoulの質問です。このスキルはリクエストが求める 具体的なことを行うか? それぞれが独立に答えられるので、すべて低く返ることもあり、 最も高いものでも 0.30 未満になった候補リストは完全に破棄されます。
RERANK_INSTRUCTIONS = (
"Exactly one of these skills is the right one to load for the user's latest "
"request. Which one? Read what each actually does, not just its name."
)
def rerank_criteria(names: tuple[str, ...], excerpt: int) -> dict[str, str]:
return {
name: f"{BY_NAME[name]['description_full']} — {BY_NAME[name]['body'][:excerpt]}"
for name in names
}
def rerank_questions(names: tuple[str, ...], excerpt: int) -> dict:
questions = {
"which": Choice(
instructions=RERANK_INSTRUCTIONS, criteria=rerank_criteria(names, excerpt)
)
}
for name in names:
questions[f"fits::{name}"] = Noul(
instructions=(
f"Does the skill '{name}' do the specific thing the user's request asks "
f"for? It is described as: {BY_NAME[name]['description_full']}"
)
)
return questions
@json_cache
def rerank(request: str, names: tuple[str, ...], excerpt: int) -> dict:
"""Request 2: the same Choice over a shortlist, plus one absolute noul per candidate."""
started = perf_counter()
response = client.system_one(
state=build_state(request),
questions=rerank_questions(names, excerpt),
model=TYPESAFE_MODEL,
)
return {
"winner": response.answers["which"].choice,
"fits": {
key.removeprefix("fits::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("fits::")
},
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
for request in DEMO:
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
print(f'"{request[:78]}"\n scored too low, nothing suggested\n')
continue
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
best = max(result["fits"].values())
verdict = result["winner"] if best >= FITS_THRESHOLD else "nothing fits"
print(f'"{request[:78]}"')
print(f" was {shortlist[0]} -> {verdict} ({result['seconds']}s)")
for name in shortlist:
print(f" fits {result['fits'][name]:.2f} {name}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
was apple-notes -> apple-notes (0.12s)
fits 0.60 apple-notes
fits 0.54 computer-use
fits 0.01 concept-diagrams
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
was powerpoint -> pptx-author (0.09s)
fits 0.73 powerpoint
fits 0.38 pptx-author
fits 0.02 chroma
"Post this announcement to my Mastodon account."
was xurl -> xurl (0.09s)
fits 0.56 xurl
fits 0.38 computer-use
fits 0.05 openhands
2 つの .pptx スキルは、それぞれが自身のテキストを持ち込むと分かれます。デックの
リクエストは作成スキルに反転します。
ここで fits の noul と Choice が食い違います。noul は編集スキルを高く採点し、Choice は
作成スキルを選びます。両者は別のことを決めています。Choice は どの スキルかを決め、
noul はそもそも何かを言うべきか どうか を決めます。
Mastodon のリクエストは両方のチェックを生き延びます。最も高い fits noul が 0.30 を
超えるので、レシピは Mastodon についてのリクエストに X のスキルを提案します。この種の
リクエストのほとんどは捉えられます。2 回目のパスは、広い順位付けが渡したものしか
却下できず、ここでは 3 つの惜しい外れでした。
以下の関数がレシピのすべてです。2 回のリクエストと 2 つのしきい値で、返ってくるスキル名は 最大 1 つです。
自分の名簿に向けるには、hermes_roster.json を置き換えます。上のすべての質問はその
ファイルから name、description、description_full、body を読み取り、他に Hermes を
知っているものはありません。
def suggest(request: str) -> tuple[str, ...]:
"""At most one skill name for a request, or () for "nothing here applies"."""
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
return ()
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
if max(result["fits"].values()) < FITS_THRESHOLD:
return ()
return (result["winner"],)
def suggestion_block(names: tuple[str, ...]) -> str:
"""What gets appended after the roster, in the suggestion.
This string is a measured input rather than prose: it goes to the agent, so it is part
of every graded turn's cache key. Editing a word here silently invalidates the shipped
results and costs a live re-run to restore them.
"""
body = (
f"Relevant to the current request: {', '.join(names)}. Ignore this if it does not "
"fit what the user actually asked for."
if names
else "No skill in the roster appears relevant to this request."
)
return f"\n\n<skill_relevance>\n{body}\n</skill_relevance>"
print(suggestion_block(suggest(DEMO[1])))
print(suggestion_block(suggest(DEMO[2])))
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
<skill_relevance>
Relevant to the current request: xurl. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
ステップ 5:提案を測定する
488 件のリクエストそれぞれが、エージェントに 3 回、それぞれ 1 つの測定ターンとして 渡されます。実行の違いは、エージェントに何を伝えるかだけです。
| システムプロンプトに入るもの | |
|---|---|
| エージェント単体 | なし |
| 提案を受けたエージェント | suggest() が返したもの |
| 正解を渡されたエージェント | 正解のスキル名、なければ「何も当てはまらない」 |
3 つ目は実現可能ではありません。他の 2 つを測る際の上限です。
その提案の文言は 2 つの役割を果たします。提案は無視してよいと述べるのは、強く押し出すと 誤った提案にも従わせてしまい、誤った提案は何もないより悪いからです。また、提案するものが ないターンでも、そう述べる 1 文を依然として送ります。何も送らなければ、名簿自身の 「読み込む側に倒す」という instruction が対抗されないままになります。
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool: # up to 488 x 2 TypeSafe requests
suggested = dict(zip(texts, pool.map(suggest, texts)))
WIDE = {text: rank_wide(text) for text in texts} # all cache hits now; reused below
arms = {
"baseline": {},
"TypeSafe": {
request["text"]: suggestion_block(suggested[request["text"]])
for request in REQUESTS
},
"oracle": {
request["text"]: suggestion_block((request["gold"],) if request["gold"] else ())
for request in REQUESTS
},
}
scores = {
arm: summarise(run_arm(arm, suggestions)) for arm, suggestions in arms.items()
}
print(f"{'run':<10}{'wrong loads':>13}{'needless loads':>16}")
for arm, row in scores.items():
print(f"{arm:<10}{row['wrong_load']:>13.1%}{row['needless_load']:>16.1%}")
def fewer(metric: str) -> str:
"""The plain ratio between the two arms' error rates."""
return f"{scores['baseline'][metric] / scores['TypeSafe'][metric]:.1f}x fewer"
print(
f"\nbaseline -> TypeSafe: {fewer('wrong_load')} wrong loads, "
f"{fewer('needless_load')} needless ones"
)
run wrong loads needless loads
baseline 16.8% 9.8%
TypeSafe 7.3% 4.0%
oracle 2.5% 1.2%
baseline -> TypeSafe: 2.3x fewer wrong loads, 2.4x fewer needless ones
moved = [
(
baseline[p["text"]]["loaded"][:1] == [p["gold"]],
run_turn(AGENT_MODEL, "TypeSafe", p["text"], arms["TypeSafe"][p["text"]])[
"loaded"
][:1]
== [p["gold"]],
)
for p in POSITIVES
]
print(
f"of {len(POSITIVES)} covered requests: {sum(not b and a for b, a in moved)} the suggestion "
f"fixed, {sum(b and not a for b, a in moved)} it broke"
)
of 315 covered requests: 37 the suggestion fixed, 7 it broke
提案は壊すものよりはるかに多くを直しますが、エージェントが単体で正しくできていたものを 壊すこともあります。自信のある誤った提案は、提案がまったくないより説得力があり、それが ターンの前に提案を置く代償です。
SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"
ARM_COLOR = {"baseline": BLUE, "TypeSafe": ORANGE, "oracle": MUTED}
def style(ax):
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
panels = [
("wrong_load", f"wrong loads\n{len(POSITIVES)} covered requests"),
("needless_load", f"needless loads\n{len(NEGATIVES)} uncovered requests"),
]
names = list(scores)
fig, axes = plt.subplots(1, 2, figsize=(8.4, 3.6), facecolor=SURFACE)
for ax, (metric, title) in zip(axes, panels):
style(ax)
ax.grid(axis="y", color=GRID, linewidth=0.8)
values = [scores[arm][metric] for arm in names]
bars = ax.bar(
names,
values,
0.58,
color=[ARM_COLOR[arm] for arm in names],
# the oracle is a ceiling, not a competitor: gray, and hatched so it never depends
# on colour alone
hatch=["", "", "///"],
edgecolor=SURFACE,
linewidth=1.2,
)
ax.bar_label(
bars,
labels=[f"{v:.1%}" for v in values],
padding=3,
color=INK2,
fontsize=9,
)
ax.set_title(title, loc="left", color=INK2, fontsize=9.5)
ax.set_ylim(0, max(values) * 1.28)
ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))
ax.set_ylabel("% of those requests - lower is better", color=INK2, fontsize=9)
fig.suptitle(
f"Hermes' {len(ROSTER)}-skill roster, {len(REQUESTS)} requests, {AGENT_MODEL}",
x=0.02,
ha="left",
color=INK,
fontsize=11,
)
fig.tight_layout()
display(fig)
plt.close(fig)
結果が示すもの
- 誤った読み込みは 16.8% から 7.3% に、不要な読み込みは 9.8% から 4.0% に下がりました。 これは、切り詰められたインデックスからの推測と、正解を渡されることの差のほとんどです。
- エージェントが単体で正しくできていたリクエストのいくつかは、提案が付くと誤って 返ってきます。件数は上にあります。
自分のエージェントが大きな名簿を持つときは、この形をコピーしてください。すべてに対する 安価な順位付けと、2 つか 3 つへの詳しい検討です。どちらのステップも何も得られずに 終わることがあります。
playground で開く
ステップ 4 のデックのリクエスト用に、playground のリンクを組み立てます。各候補の完全な 説明と body の抜粋を criteria として使います。
demo_shortlist = tuple(name for name, _ in rank_wide(DEMO[1])["ranked"][:SHORTLIST])
playground_link = make_playground_link(
build_state(DEMO[1]),
rerank_questions(demo_shortlist, EXCERPT_CHARS),
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open the shortlist + questions in the TypeSafe playground]({playground_link})"
)
)
TypeSafe playground で候補リストと質問を開く →
次に読むもの
同じ形は他の場所にも現れます。スキルではなくハンドラへのルーティングなら Intent Routing、2 つのしきい値の選び方なら Confidence、すべての質問を 1 回のリクエストに入れるなら Speculative Fan-Out です。