Skill-Vorschlag
Wählt für einen Agenten-Turn höchstens einen Skill aus den 182 in Nous Researchs Hermes-Katalog aus, mit zwei TypeSafe-Anfragen zum Reihen und Nachprüfen der besten Kandidaten.
Agenten wählen Skills, indem sie sie alle abschneiden und in die Systemnachricht laden, was die Kosten erhöht, die Skill-Auswahl verschlechtert und für den Rest der Sitzung Kontextverfall auslöst. Wir begegnen dem mit zwei TypeSafe-Anfragen pro Turn, einer zum Reihen der Skills und einer zum Überprüfen der Wahl, und reduzieren falsche Skill-Ladevorgänge um mehr als die Hälfte.
Ein Agent mit einem großen Skill-Verzeichnis trifft seine Wahl auf fast keiner Grundlage. Das
Verzeichnis erreicht ihn als Index: eine Zeile pro Skill, wobei die Beschreibung abgeschnitten
ist, damit der Volltext nicht die Konversation verdrängt. Hermes, das hier verwendete
Agenten-Harness, kürzt sie standardmäßig auf 60 Zeichen. Bei dieser Breite liest sich der Skill,
der .pptx-Dateien bearbeitet, fast genauso wie der, der sie erstellt. Bitte um ein
Pitch-Deck, und der Agent lädt womöglich den falschen. In einem Turn, in dem gar kein Skill
passt, lädt er trotzdem einen, weil eine Liste von Namen zum Raten einlädt.
Dieses Cookbook lässt die Beschreibungen unangetastet und setzt stattdessen auf progressive Offenlegung: Es liest alle 182 Skills günstig und dann drei davon im Detail. Zwei TypeSafe-Anfragen stehen vor der Entscheidung, welcher Skill geladen werden soll, falls überhaupt einer. Die erste reiht jeden Skill im Verzeichnis gegen den Turn des Nutzers und beantwortet, ob der Turn überhaupt einen Skill braucht. Die zweite liest nur die obersten drei erneut, jetzt mit der vollständigen Beschreibung jedes Skills und dem Anfang seiner Anweisungen, und darf alle ablehnen.
Der Name des Gewinners kommt in eine zusätzliche Zeile des System-Prompts des Agenten für diesen Turn:
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user
actually asked for.
</skill_relevance>
Der Agent behält seinen vollständigen Index und sein eigenes Urteil, und diese eine Zeile sagt
ihm nur, welchen Eintrag er zuerst ansehen soll. Das Verzeichnis selbst ändert sich nie, also
hält jedes Prefix-Caching darauf weiter. Über 488 Anfragen gegen claude-haiku-4-5-20251001,
mit Skills aus dem Hermes-Verzeichnis:
| lädt den falschen Skill | lädt einen, wenn keiner passt | |
|---|---|---|
| Agent allein, nur mit seinem Verzeichnis | 16.8% | 9.8% |
| Agent mit einem TypeSafe-Vorschlag | 7.3% | 4.0% |
| Agent mit der richtigen Antwort | 2.5% | 1.2% |
Die dritte Zeile zeigt, dass der Boden für Fehler nicht null ist, denn ein Agent, dem der richtige Skill gegeben wird, lädt ihn trotzdem nicht immer, und keine Auswahlmethode, so gut sie auch sei, kommt darüber hinweg.
Am Ende hast du eine suggest()-Funktion, die höchstens einen Skill-Namen zurückgibt, einen
suggestion_block(), der ihn für den System-Prompt verpackt, und das Harness, das die obige
Tabelle erzeugt hat, bereit, auf dein eigenes Verzeichnis gerichtet zu werden.
flowchart LR
subgraph C1["Call 1 - skim all 182 skills"]
direction TB
Q1["<b>Choice:</b> which skill fits?<br/><i>all 182, one line each</i>"]
N1["<b>Nouls:</b> need a skill at all?<br/>· act on their stuff?<br/>· follow written steps?<br/>· or just talk?"]
%% invisible link: without an edge these two share a rank, which in a TB
%% subgraph puts them side by side instead of stacked
Q1 ~~~ N1
end
subgraph C2["Call 2 - read those 3 properly"]
direction TB
Q2["<b>Choice:</b> which of the 3?<br/><i>with real detail now</i>"]
N2["<b>Nouls:</b> does each one<br/>really do it?"]
Q2 ~~~ N2
end
REQ["the request"] --> C1
C1 -->|"top 3"| C2
C1 -->|"nothing<br/>applies"| STOP["suggest<br/>nothing"]
C2 -->|"none fit"| STOP
C2 -->|"a winner"| OUT["suggest<br/>the winner"]
Einrichtung
- Installiere den TypeSafe-Client, den Anthropic-Client und die gemeinsamen Cookbook-Helfer.
- Setze einen TypeSafe-API-Schlüssel und einen Anthropic-Schlüssel für den Agenten, der gemessen wird.
pip install anthropic matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
export TYPESAFE_API_KEY=your-key-here
export ANTHROPIC_API_KEY=your-key-here
Hinweis: Die Codeblöcke unten sind ein Skript, in Reihenfolge. Um mitzumachen, packe sie in eine einzelne Datei in der gezeigten Reihenfolge.
Ergebnisse cachen
JsonCache speichert das Ergebnis jedes Aufrufs, mit seinen Eingaben als Schlüssel, sodass ein
erneuter Lauf die Zahlen unten erneut abspielt, statt eine der beiden APIs aufzurufen. Lösche
json_cache.json, um live zu laufen. Der veröffentlichte Lauf verwendete jev-1.12 und
claude-haiku-4-5-20251001, gerendert am 2026-07-31.
import json
import os
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from time import perf_counter
import anthropic
import matplotlib
import matplotlib.pyplot as plt
from matplotlib.ticker import PercentFormatter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient
matplotlib.use("Agg") # headless render
TYPESAFE_MODEL = "jev-1.12"
AGENT_MODEL = (
"claude-haiku-4-5-20251001" # the agent under test, pinned so scores are stable
)
SHORTLIST = 3 # candidates carried from the first request into the second
EXCERPT_CHARS = (
700 # SKILL.md characters each candidate brings; the roster file stores 1600
)
GATE_THRESHOLD = (
0.30 # mean of the three request nouls, below which nothing is suggested
)
FITS_THRESHOLD = (
0.30 # a shortlist whose best "does this fit" noul is under this is dropped
)
WORKERS = 8 # small pool: enough to keep a live run to minutes, gentle on rate limits
assert EXCERPT_CHARS <= 1600, (
"the shipped roster file stores 1600 body characters per skill"
)
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
agent = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "cache-only"))
json_cache = JsonCache(Path("json_cache.json"))
Schritt 1: Das Verzeichnis laden
hermes_roster.json enthält die 182 Skills von
NousResearch/hermes-agent (MIT) an einem
fest angehefteten Commit. Jeder Eintrag enthält Name und Kategorie eines Skills, die
Beschreibung, wie der Index sie zeigt, die vollständige Beschreibung und den Anfang seiner
SKILL.md.
Der Index unten und die Anweisungen darüber im Prompt sind aus Hermes kopiert.
ROSTER = json.loads(Path("hermes_roster.json").read_text(encoding="utf-8"))
BY_NAME = {skill["name"]: skill for skill in ROSTER}
# Verbatim from hermes-agent agent/prompt_builder.py:build_skills_system_prompt.
PREAMBLE = (
"## Skills (mandatory)\n"
"Before replying, scan the skills below. If a skill matches or is even partially relevant "
"to your task, you MUST load it with skill_view(name) and follow its instructions. "
"Err on the side of loading — it is always better to have context you don't need "
"than to miss critical steps, pitfalls, or established workflows. "
"Skills contain specialized knowledge — API endpoints, tool-specific commands, "
"and proven workflows that outperform general-purpose approaches. Load the skill "
"even if you think you could handle the task with basic tools like web_search or terminal. "
"Skills also encode the user's preferred approach, conventions, and quality standards "
"for tasks like code review, planning, and testing — load them even for tasks you "
"already know how to do, because the skill defines how it should be done here.\n"
"Whenever the user asks you to configure, set up, install, enable, disable, modify, "
"or troubleshoot Hermes Agent itself — its CLI, config, models, providers, tools, "
"skills, voice, gateway, plugins, or any feature — load the `hermes-agent` skill "
"first. It has the actual commands (e.g. `hermes config set …`, `hermes tools`, "
"`hermes setup`) so you don't have to guess or invent workarounds.\n"
"If a skill has issues, fix it with skill_manage(action='patch').\n"
"After difficult/iterative tasks, offer to save as a skill. "
"If a skill you loaded was missing steps, had wrong commands, or needed "
"pitfalls you discovered, update it before finishing.\n"
"\n"
)
FOOTER = "\n\nOnly proceed without loading a skill if genuinely none are relevant to the task."
IDENTITY = (
"You are Hermes, a capable AI assistant with access to tools and a library "
"of skills. You help the user with coding, research, and everyday tasks.\n\n"
)
def render_index() -> str:
"""The body of <available_skills>: skills grouped by category, both sorted by name."""
by_category = defaultdict(list)
for skill in ROSTER:
by_category[skill["category"]].append(skill)
lines = []
for category in sorted(by_category):
lines.append(f" {category}:")
for skill in sorted(by_category[category], key=lambda s: s["name"]):
lines.append(f" - {skill['name']}: {skill['description']}")
return "\n".join(lines)
CATALOG_PROMPT = (
IDENTITY
+ PREAMBLE
+ "<available_skills>\n"
+ render_index()
+ "\n</available_skills>"
+ FOOTER
)
widths = [len(skill["description"]) for skill in ROSTER]
print(f"{len(ROSTER)} skills in {len({s['category'] for s in ROSTER})} categories")
print(f"roster prompt: {len(CATALOG_PROMPT):,} characters")
print(
f"index description: {sum(widths) / len(widths):.0f} characters on average, "
f"{max(widths)} at most"
)
print("\none category, as the agent reads it:")
index_lines = render_index().splitlines()
start = index_lines.index(" apple:")
end = next(
i
for i in range(start + 1, len(index_lines))
if not index_lines[i].startswith(" ")
)
print("\n".join(index_lines[start:end]))
182 skills in 33 categories
roster prompt: 16,089 characters
index description: 54 characters on average, 60 at most
one category, as the agent reads it:
apple:
- apple-notes: Manage Apple Notes via memo CLI: create, search, edit.
- apple-reminders: Apple Reminders via remindctl: add, list, complete.
- findmy: Track Apple devices/AirTags via FindMy.app on macOS.
- imessage: Send and receive iMessages/SMS via the imsg CLI on macOS.
Schritt 2: Den Agenten allein bewerten
requests.json enthält 488 Single-Turn-Anfragen, 315 davon von genau einem Skill abgedeckt
und die anderen 173 von keinem.
Die abgedeckten Anfragen wurden von Claude Sonnet 5 aus der jeweiligen SKILL.md jedes
Skills geschrieben, also sind die Labels vertrauenswürdig und die Anfragen leichter als die,
die Nutzer senden.
Die 173 nicht abgedeckten wurden alle geschrieben, um Raten zu bestrafen: 85 Alltagsanfragen, 42 technische Fragen, die kein Skill bedient (erkläre, was eine Monade ist), und 46, die nach etwas Bestimmtem fragen, für das das Verzeichnis keinen Skill hat, etwa poste das auf Mastodon in einem Verzeichnis, das X und sonst nichts abdeckt.
Die Bewertung liest nur die erste Antwort des Agenten. Beide Zahlen sind Fehlerraten, also ist bei jeder ein niedrigerer Wert besser:
- wrong load: der Anteil der abgedeckten Anfragen, bei denen der erste
skill_view-Aufruf nicht der abdeckende Skill war. Ein Turn, der gar nichts geladen hat, zählt als Fehlgriff. - needless load: der Anteil der nicht abgedeckten Anfragen, bei denen der Agent
überhaupt
skill_viewaufgerufen hat.
REQUESTS = json.loads(Path("requests.json").read_text(encoding="utf-8"))
POSITIVES = [p for p in REQUESTS if p["gold"]]
NEGATIVES = [p for p in REQUESTS if not p["gold"]]
print(
f"{len(REQUESTS)} requests: {len(POSITIVES)} covered by a skill "
f"({len({p['gold'] for p in POSITIVES})} distinct skills), {len(NEGATIVES)} covered by none"
)
print(f"\ncovered [{POSITIVES[0]['gold']}] {POSITIVES[0]['text']}")
print(f"uncovered {NEGATIVES[0]['text']}")
488 requests: 315 covered by a skill (171 distinct skills), 173 covered by none
covered [1password] I've got a config.yaml with `{{ op://app-prod/db/password }}` placeholders in it — can you set up my project to pull the real values in at runtime instead of hardcoding them?
uncovered Add these three cards to our Trello backlog.
Der Vorschlag kommt in einen eigenen Block des System-Prompts, nach dem Verzeichnis statt darin, damit der Verzeichnistext in jedem Turn identisch bleibt und das Prefix-Caching erhalten bleibt.
Der Agent hat einen minimalen Satz an Tools, darunter skill_view, um einen Skill über einen
Freitext-Namen zu laden. Der Name muss exakt zum Skill passen, damit der Ladevorgang korrekt
ist.
# Verbatim from hermes-agent tools/skills_tool.py:SKILL_VIEW_SCHEMA.
SKILL_VIEW_DESCRIPTION = (
"Skills allow for loading information about specific tasks and workflows, as "
"well as scripts and templates. Load a skill's full content or access its "
"linked files (references, templates, scripts). First call returns SKILL.md "
"content plus a 'linked_files' dict showing available references/templates/"
"scripts. To access those, call again with file_path parameter."
)
TOOLS = [
{
"name": "skill_view",
"description": SKILL_VIEW_DESCRIPTION,
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string", "description": "The skill name."}
},
"required": ["name"],
},
},
{
"name": "terminal",
"description": "Run a shell command on the user's machine and return its output.",
"input_schema": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"],
},
},
{
"name": "read_file",
"description": "Read a file from the user's filesystem.",
"input_schema": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"],
},
},
{
"name": "web_search",
"description": "Search the web and return result snippets.",
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
]
@json_cache
def run_turn(model: str, arm: str, request: str, suggestion: str) -> dict:
"""One measured turn. ``arm`` is in the key so each arm samples independently."""
system = [
{"type": "text", "text": CATALOG_PROMPT, "cache_control": {"type": "ephemeral"}}
]
if suggestion:
system.append({"type": "text", "text": suggestion}) # after the breakpoint
response = agent.messages.create(
model=model,
max_tokens=1024,
system=system,
tools=TOOLS,
messages=[{"role": "user", "content": request}],
)
usage = response.usage
return {
"loaded": [
str(block.input.get("name", ""))
for block in response.content
if block.type == "tool_use" and block.name == "skill_view"
],
"input_tokens": usage.input_tokens or 0,
"output_tokens": usage.output_tokens or 0,
}
def summarise(turns: dict[str, dict]) -> dict[str, float]:
"""Two failure rates: wrong loads on covered requests, needless ones on uncovered."""
hits = [turns[p["text"]]["loaded"][:1] == [p["gold"]] for p in POSITIVES]
over = [bool(turns[p["text"]]["loaded"]) for p in NEGATIVES]
return {
# both metrics are errors, so the two columns read the same direction
"wrong_load": 1 - sum(hits) / len(hits),
"needless_load": sum(over) / len(over),
}
def run_arm(arm: str, suggestions: dict[str, str]) -> dict[str, dict]:
"""One measured turn per request, in a small pool. 488 calls."""
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
turns = pool.map(
lambda t: run_turn(AGENT_MODEL, arm, t, suggestions.get(t, "")), texts
)
return dict(zip(texts, turns))
Der Agent läuft zuerst mit nichts als seinem Verzeichnis, so wie er heute arbeitet. Seine beiden Fehlerraten sind die Baseline, an der der Rest des Cookbooks misst.
baseline = run_arm("baseline", {})
base_scores = summarise(baseline)
print(
f"wrong loads {base_scores['wrong_load']:.1%} ({len(POSITIVES)} covered requests)"
)
print(
f"needless loads {base_scores['needless_load']:.1%} ({len(NEGATIVES)} uncovered requests)"
)
# where the wrong loads land: a neighbour of the right skill, or somewhere unrelated?
misses = [
(p["gold"], baseline[p["text"]]["loaded"][0])
for p in POSITIVES
if baseline[p["text"]]["loaded"] and baseline[p["text"]]["loaded"][0] != p["gold"]
]
same_category = sum(
1
for gold, got in misses
if got in BY_NAME and BY_NAME[got]["category"] == BY_NAME[gold]["category"]
)
print(
f"\nof {len(misses)} wrong first picks, {same_category} came from the right skill's own "
f"category"
)
wrong loads 16.8% (315 covered requests)
needless loads 9.8% (173 uncovered requests)
of 36 wrong first picks, 10 came from the right skill's own category
Falsche Ladevorgänge landen weit häufiger in der eigenen Kategorie des richtigen Skills, als der Zufall sie dort platzieren würde, also liegt die Schwierigkeit darin, ein paar zum Verwechseln ähnliche Skills auseinanderzuhalten. Der Agent schaut bereits ungefähr an der richtigen Stelle.
Schritt 3: Das ganze Verzeichnis reihen
Eine Anfrage trägt zwei Arten von Frage:
whichist eineChoice-Frage über alle 182 Skill-Namen, mit der Index-Beschreibung als Kriterien jeder Option (derselbe Text, den der Agent selbst bekommt). Ihre Wahrscheinlichkeiten sind die Reihung.- drei
Noul-Fragen zur Anfrage, unten ausgegeben, die jeweils auf andere Weise fragen, ob eine Aktion gewünscht ist statt einer Erklärung.prose_sufficeszählt andersherum. Ihr Mittel entscheidet, ob überhaupt etwas vorgeschlagen wird, und unter 0.30 wird nichts vorgeschlagen.
Beide gehen in einer Anfrage raus, also kosten Reihung und Prüfung einen Round Trip.
Formuliere diese drei so, dass sie fragen, ob eine Aktion gewünscht ist. Eine Frage zum Thema wird erkläre, was eine Monade ist nicht von einer Anfrage trennen, die einen Skill braucht, da beides Software ist.
Eine Choice-Frage hält ein Verzeichnis dieser Größe bequem. Ein paar Mal größer, und du
würdest
es in Blöcke aufteilen und jeden einzeln reihen, dann diesen selben Shortlist-Schritt über die
Gewinner laufen lassen.
CHOICE_INSTRUCTIONS = (
"Which of these skills, if any, is the right one to load to help with the "
"user's latest request?"
)
GATE_QUESTIONS = {
"acts_on_user_system": (
"Is the assistant being asked to act on the user's files, accounts, devices, "
"or online services, rather than only to explain or advise?"
),
"would_follow_documented_procedure": (
"Would a careful expert answering this consult a specific documented procedure "
"or set of commands, rather than answering from general understanding?"
),
"prose_suffices": (
"Could a knowledgeable generalist fully satisfy this request in prose, with "
"no tools, no documentation, and no access to the user's files or accounts?"
),
}
INVERTED = {"prose_suffices"} # a yes here points away from needing a skill
def build_state(request: str) -> dict:
return {"request": request, "recent_context": ""}
@json_cache
def rank_wide(request: str) -> dict:
"""Request 1: rank all 182 skills, and score the request for whether a skill applies."""
questions = {
"which": Choice(
instructions=CHOICE_INSTRUCTIONS,
criteria={skill["name"]: skill["description"] for skill in ROSTER},
)
}
for key, text in GATE_QUESTIONS.items():
questions[f"gate::{key}"] = Noul(instructions=text)
started = perf_counter()
response = client.system_one(
state=build_state(request), questions=questions, model=TYPESAFE_MODEL
)
ranked = sorted(
response.answers["which"].probabilities.items(), key=lambda kv: -kv[1]
)
values = {
key.removeprefix("gate::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("gate::")
}
oriented = [(1.0 - v) if k in INVERTED else v for k, v in values.items()]
return {
"ranked": ranked[
:12
], # more than any shortlist needs, and keeps the cache small
"gate": sum(oriented) / len(oriented),
"values": values,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
DEMO = [
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so it syncs"
" to my phone? Just write it up in whatever editor pops up.",
"Can you put together a pitch deck skeleton (cover, situation overview, comps, precedent"
" transactions, DCF, LBO) as a .pptx, using our firm-template.pptx for branding and"
" footnoting each valuation number back to the cell it came from in the model?",
"Post this announcement to my Mastodon account.",
]
for request in DEMO:
wide = rank_wide(request)
verdict = "suggest" if wide["gate"] >= GATE_THRESHOLD else "stay quiet"
print(f'"{request[:78]}"')
print(f" needs a skill {wide['gate']:.2f} -> {verdict} ({wide['seconds']}s)")
for name, probability in wide["ranked"][:SHORTLIST]:
print(f" {probability:.3f} {name:<38}{BY_NAME[name]['description']}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
needs a skill 0.75 -> suggest (0.31s)
0.990 apple-notes Manage Apple Notes via memo CLI: create, search, edit.
0.010 computer-use Drive the user's desktop in the background — clicking, ty...
0.000 concept-diagrams Generate flat, minimal educational SVG visuals as HTML.
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
needs a skill 0.76 -> suggest (0.16s)
0.700 powerpoint Create, read, edit .pptx decks, slides, notes, templates.
0.300 pptx-author Build PowerPoint decks headless with python-pptx.
0.000 chroma Embedding database for RAG and semantic search.
"Post this announcement to my Mastodon account."
needs a skill 0.78 -> suggest (0.16s)
0.550 xurl X/Twitter via xurl CLI: raw post search, posting, DM, media.
0.140 computer-use Drive the user's desktop in the background — clicking, ty...
0.080 openhands Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).
Die Notes.app-Anfrage ist eindeutig, und ihre oberste Option ist die richtige. Nichts, was eine Reihung leisten kann, rettet die Mastodon-Anfrage: Die drei Fragen sagen, dass ein Skill gewünscht ist, weil das Posten auf ein Konto eine Aktion ist, und mit einem Skill für das Posten auf X und nichts für Mastodon gewinnt sowieso der nächstbeste Skill.
Bleibt das Deck. Beide Spitzenreiter sind .pptx-Skills, und bei 60 Zeichen setzt die breite
Choice-Frage den bearbeitenden Skill vor den erstellenden, für eine Anfrage über
die Erstellung eines Decks.
Schritt 4: Die obersten drei neu reihen
Drei Optionen lassen Raum für die vollständige Beschreibung plus den Anfang der jeweiligen
SKILL.md jedes Skills, also legt die zweite Anfrage dieselbe Frage besseren Belegen vor:
whichist eineChoice-Frage über die Shortlist, mit diesem längeren Text als Kriterien jeder Option.fits::{name}ist eineNoul-Frage pro Kandidat: Tut dieser Skill das Konkrete, wonach die Anfrage fragt? Jede wird für sich beantwortet, also können alle niedrig zurückkommen, und eine Shortlist, deren höchste unter 0.30 landet, wird ganz verworfen.
RERANK_INSTRUCTIONS = (
"Exactly one of these skills is the right one to load for the user's latest "
"request. Which one? Read what each actually does, not just its name."
)
def rerank_criteria(names: tuple[str, ...], excerpt: int) -> dict[str, str]:
return {
name: f"{BY_NAME[name]['description_full']} — {BY_NAME[name]['body'][:excerpt]}"
for name in names
}
def rerank_questions(names: tuple[str, ...], excerpt: int) -> dict:
questions = {
"which": Choice(
instructions=RERANK_INSTRUCTIONS, criteria=rerank_criteria(names, excerpt)
)
}
for name in names:
questions[f"fits::{name}"] = Noul(
instructions=(
f"Does the skill '{name}' do the specific thing the user's request asks "
f"for? It is described as: {BY_NAME[name]['description_full']}"
)
)
return questions
@json_cache
def rerank(request: str, names: tuple[str, ...], excerpt: int) -> dict:
"""Request 2: the same Choice over a shortlist, plus one absolute noul per candidate."""
started = perf_counter()
response = client.system_one(
state=build_state(request),
questions=rerank_questions(names, excerpt),
model=TYPESAFE_MODEL,
)
return {
"winner": response.answers["which"].choice,
"fits": {
key.removeprefix("fits::"): answer.noul
for key, answer in response.answers.items()
if key.startswith("fits::")
},
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
for request in DEMO:
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
print(f'"{request[:78]}"\n scored too low, nothing suggested\n')
continue
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
best = max(result["fits"].values())
verdict = result["winner"] if best >= FITS_THRESHOLD else "nothing fits"
print(f'"{request[:78]}"')
print(f" was {shortlist[0]} -> {verdict} ({result['seconds']}s)")
for name in shortlist:
print(f" fits {result['fits'][name]:.2f} {name}")
print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
was apple-notes -> apple-notes (0.12s)
fits 0.60 apple-notes
fits 0.54 computer-use
fits 0.01 concept-diagrams
"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
was powerpoint -> pptx-author (0.09s)
fits 0.73 powerpoint
fits 0.38 pptx-author
fits 0.02 chroma
"Post this announcement to my Mastodon account."
was xurl -> xurl (0.09s)
fits 0.56 xurl
fits 0.38 computer-use
fits 0.05 openhands
Die beiden .pptx-Skills trennen sich, sobald jeder seinen eigenen Text mitbringt: Die
Deck-Anfrage kippt zum erstellenden Skill.
Die fits-Nouls und die Choice widersprechen sich dort: Die Nouls bewerten den bearbeitenden
Skill höher, während die Choice den erstellenden wählt. Sie entscheiden Unterschiedliches. Die
Choice klärt, welcher Skill, und die Nouls klären, ob überhaupt etwas gesagt werden soll.
Die Mastodon-Anfrage übersteht beide Prüfungen: Ihr bester fits-Noul landet über 0.30, also
schlägt das Rezept den X-Skill für eine Anfrage über Mastodon vor. Die meisten ähnlichen
Anfragen werden abgefangen. Der zweite Durchgang kann nur ablehnen, was die breite Reihung ihm
übergibt, und hier waren das drei Beinahe-Treffer.
Die Funktion unten ist das ganze Rezept: zwei Anfragen und zwei Schwellenwerte, wobei höchstens ein Skill-Name zurückkommt.
Um es auf dein eigenes Verzeichnis zu richten, ersetze hermes_roster.json. Jede Frage oben
liest name, description, description_full und body aus dieser Datei, und sonst weiß
nichts von Hermes.
def suggest(request: str) -> tuple[str, ...]:
"""At most one skill name for a request, or () for "nothing here applies"."""
wide = rank_wide(request)
if wide["gate"] < GATE_THRESHOLD:
return ()
shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
result = rerank(request, shortlist, EXCERPT_CHARS)
if max(result["fits"].values()) < FITS_THRESHOLD:
return ()
return (result["winner"],)
def suggestion_block(names: tuple[str, ...]) -> str:
"""What gets appended after the roster, in the suggestion.
This string is a measured input rather than prose: it goes to the agent, so it is part
of every graded turn's cache key. Editing a word here silently invalidates the shipped
results and costs a live re-run to restore them.
"""
body = (
f"Relevant to the current request: {', '.join(names)}. Ignore this if it does not "
"fit what the user actually asked for."
if names
else "No skill in the roster appears relevant to this request."
)
return f"\n\n<skill_relevance>\n{body}\n</skill_relevance>"
print(suggestion_block(suggest(DEMO[1])))
print(suggestion_block(suggest(DEMO[2])))
<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
<skill_relevance>
Relevant to the current request: xurl. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>
Schritt 5: Den Vorschlag messen
Jede der 488 Anfragen geht dreimal an den Agenten, jeweils ein gemessener Turn. Die Läufe unterscheiden sich nur darin, was dem Agenten gesagt wird:
| was in den System-Prompt kommt | |
|---|---|
| Agent allein | nichts |
| Agent mit Vorschlag | was auch immer suggest() zurückgab |
| Agent mit der Antwort | der Name des abdeckenden Skills oder „nichts passt“, wenn es keinen gibt |
Der dritte ist nicht erreichbar; er ist die Obergrenze, an der die anderen beiden gemessen werden.
Die Formulierung dieses Vorschlags erfüllt zwei Aufgaben. Sie sagt, dass der Vorschlag ignoriert werden darf, weil stärkeres Drängen auch bei falschen Vorschlägen Befolgung erzwingt und ein falscher schlimmer ist als keiner. Und ein Turn ohne etwas vorzuschlagen sendet trotzdem einen Satz, der das sagt; gar nichts zu senden würde die eigene Anweisung des Verzeichnisses, „im Zweifel lieber laden“, unopponiert lassen.
texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool: # up to 488 x 2 TypeSafe requests
suggested = dict(zip(texts, pool.map(suggest, texts)))
WIDE = {text: rank_wide(text) for text in texts} # all cache hits now; reused below
arms = {
"baseline": {},
"TypeSafe": {
request["text"]: suggestion_block(suggested[request["text"]])
for request in REQUESTS
},
"oracle": {
request["text"]: suggestion_block((request["gold"],) if request["gold"] else ())
for request in REQUESTS
},
}
scores = {
arm: summarise(run_arm(arm, suggestions)) for arm, suggestions in arms.items()
}
print(f"{'run':<10}{'wrong loads':>13}{'needless loads':>16}")
for arm, row in scores.items():
print(f"{arm:<10}{row['wrong_load']:>13.1%}{row['needless_load']:>16.1%}")
def fewer(metric: str) -> str:
"""The plain ratio between the two arms' error rates."""
return f"{scores['baseline'][metric] / scores['TypeSafe'][metric]:.1f}x fewer"
print(
f"\nbaseline -> TypeSafe: {fewer('wrong_load')} wrong loads, "
f"{fewer('needless_load')} needless ones"
)
run wrong loads needless loads
baseline 16.8% 9.8%
TypeSafe 7.3% 4.0%
oracle 2.5% 1.2%
baseline -> TypeSafe: 2.3x fewer wrong loads, 2.4x fewer needless ones
moved = [
(
baseline[p["text"]]["loaded"][:1] == [p["gold"]],
run_turn(AGENT_MODEL, "TypeSafe", p["text"], arms["TypeSafe"][p["text"]])[
"loaded"
][:1]
== [p["gold"]],
)
for p in POSITIVES
]
print(
f"of {len(POSITIVES)} covered requests: {sum(not b and a for b, a in moved)} the suggestion "
f"fixed, {sum(b and not a for b, a in moved)} it broke"
)
of 315 covered requests: 37 the suggestion fixed, 7 it broke
Der Vorschlag repariert weit mehr Anfragen, als er kaputt macht, aber er macht einige kaputt, die der Agent allein richtig hatte. Ein selbstsicherer falscher Vorschlag ist überzeugender als gar kein Vorschlag, und das ist der Preis dafür, einen vor den Turn zu setzen.
SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"
ARM_COLOR = {"baseline": BLUE, "TypeSafe": ORANGE, "oracle": MUTED}
def style(ax):
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
panels = [
("wrong_load", f"wrong loads\n{len(POSITIVES)} covered requests"),
("needless_load", f"needless loads\n{len(NEGATIVES)} uncovered requests"),
]
names = list(scores)
fig, axes = plt.subplots(1, 2, figsize=(8.4, 3.6), facecolor=SURFACE)
for ax, (metric, title) in zip(axes, panels):
style(ax)
ax.grid(axis="y", color=GRID, linewidth=0.8)
values = [scores[arm][metric] for arm in names]
bars = ax.bar(
names,
values,
0.58,
color=[ARM_COLOR[arm] for arm in names],
# the oracle is a ceiling, not a competitor: gray, and hatched so it never depends
# on colour alone
hatch=["", "", "///"],
edgecolor=SURFACE,
linewidth=1.2,
)
ax.bar_label(
bars,
labels=[f"{v:.1%}" for v in values],
padding=3,
color=INK2,
fontsize=9,
)
ax.set_title(title, loc="left", color=INK2, fontsize=9.5)
ax.set_ylim(0, max(values) * 1.28)
ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))
ax.set_ylabel("% of those requests - lower is better", color=INK2, fontsize=9)
fig.suptitle(
f"Hermes' {len(ROSTER)}-skill roster, {len(REQUESTS)} requests, {AGENT_MODEL}",
x=0.02,
ha="left",
color=INK,
fontsize=11,
)
fig.tight_layout()
display(fig)
plt.close(fig)
Was die Ergebnisse zeigen
- Falsche Ladevorgänge fielen von 16.8% auf 7.3% und unnötige von 9.8% auf 4.0%, was der größte Teil der Lücke zwischen dem Raten aus einem abgeschnittenen Index und dem Übergeben der Antwort ist.
- Einige Anfragen, die der Agent allein richtig hatte, kommen falsch zurück, sobald ein Vorschlag anhängt. Die Zahlen stehen oben.
Übernimm diese Form, wenn einer deiner Agenten ein großes Verzeichnis trägt: eine günstige Reihung über alles, dann ein genauer Blick auf zwei oder drei. Jeder der beiden Schritte kann mit leeren Händen zurückkommen.
Im Playground öffnen
Baue einen Playground-Link für die Deck-Anfrage aus Schritt 4, mit der vollständigen Beschreibung und dem Textauszug jedes Kandidaten als dessen Kriterien.
demo_shortlist = tuple(name for name, _ in rank_wide(DEMO[1])["ranked"][:SHORTLIST])
playground_link = make_playground_link(
build_state(DEMO[1]),
rerank_questions(demo_shortlist, EXCERPT_CHARS),
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open the shortlist + questions in the TypeSafe playground]({playground_link})"
)
)
Öffne die Shortlist + Fragen im TypeSafe-Playground →
Wie es weitergeht
Dieselbe Form taucht auch anderswo auf: Intent-Routing für das Routing zu einem Handler statt zu einem Skill, Konfidenz für die Wahl der beiden Schwellenwerte und Spekulativer Fan-Out dafür, jede Frage in eine Anfrage zu packen.