문서

스킬 제안

Nous Research의 Hermes 카탈로그의 182개 스킬 중 에이전트 한 턴에 최대 하나를 고르며, 두 번의 TypeSafe 요청으로 상위 후보를 순위 매기고 재확인합니다.

에이전트는 모든 스킬을 잘라서 시스템 메시지에 전부 실어 넣는 방식으로 스킬을 고르며, 이 방식은 비용을 늘리고, 스킬 선택 성능을 떨어뜨리고, 세션의 나머지 동안 컨텍스트 부패를 유발합니다. 우리는 턴마다 두 번의 TypeSafe 요청으로 이를 해결합니다. 하나는 스킬 순위를 매기고 다른 하나는 그 선택을 검증하며, 잘못된 스킬 로드를 절반 이상 줄입니다.

스킬 로스터가 큰 에이전트는 거의 아무 정보도 없이 선택을 내립니다. 로스터는 인덱스 형태로 에이전트에 도착합니다. 스킬마다 한 줄이고, 설명은 잘려 있어 전체 텍스트가 대화를 밀어내지 않습니다. 여기서 사용하는 에이전트 harness인 Hermes는 기본적으로 이를 60자로 자릅니다. 예를 들어 그 폭에서는 .pptx 파일을 편집하는 스킬이 그것을 작성하는 스킬과 거의 똑같이 읽힙니다. 피치 덱을 요청하면 에이전트는 잘못된 것을 로드할 수 있습니다. 어떤 스킬도 전혀 맞지 않는 턴에서도 에이전트는 그래도 하나를 로드할 수 있습니다. 이름 목록이 추측을 부추기기 때문입니다.

이 cookbook은 설명을 그대로 두고 대신 점진적 공개를 사용합니다. 182개 스킬을 저렴하게 모두 읽은 다음, 그중 세 개를 자세히 읽습니다. 어떤 스킬을 로드할지 (있다면)에 대한 결정 앞에 두 번의 TypeSafe 요청이 놓입니다. 첫 번째는 로스터의 모든 스킬을 사용자의 턴과 비교해 순위를 매기고, 그 턴에 스킬이 필요한지 전혀 아닌지를 판단합니다. 두 번째는 상위 세 개만 다시 읽습니다. 이때는 각 스킬의 전체 설명과 지시문의 시작 부분을 함께 사용하며, 셋 모두를 거부할 자유도 있습니다.

승자의 이름은 해당 턴 동안 에이전트 시스템 프롬프트에 한 줄 더해집니다.

<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user
actually asked for.
</skill_relevance>

에이전트는 전체 인덱스와 자체 판단을 그대로 유지하며, 그 한 줄은 어느 항목을 먼저 볼지 알려줄 뿐입니다. 로스터 자체는 절대 바뀌지 않으므로 그에 대한 프리픽스 캐싱도 그대로 유지됩니다. Hermes 로스터의 스킬을 사용해 claude-haiku-4-5-20251001에 대해 수행한 488건의 요청에서:

잘못된 스킬 로드 맞는 스킬이 없는데도 로드
로스터만 가진 단독 에이전트 16.8% 9.8%
TypeSafe 제안이 있는 에이전트 7.3% 4.0%
정답을 받은 에이전트 2.5% 1.2%

세 번째 행은 실수의 하한이 0이 아님을 보여줍니다. 올바른 스킬을 받은 에이전트도 항상 그것을 로드하지는 않으며, 아무리 좋은 선택 방법도 이 점을 넘어서지 못합니다.

결국 최대 하나의 스킬 이름을 반환하는 suggest() 함수, 그것을 시스템 프롬프트용으로 감싸는 suggestion_block(), 그리고 위 표를 만들어 낸 harness를 얻게 됩니다. 자신의 로스터를 향하도록 바로 돌릴 수 있습니다.

flowchart LR
    subgraph C1["Call 1 - skim all 182 skills"]
        direction TB
        Q1["<b>Choice:</b> which skill fits?<br/><i>all 182, one line each</i>"]
        N1["<b>Nouls:</b> need a skill at all?<br/>· act on their stuff?<br/>· follow written steps?<br/>· or just talk?"]
        %% invisible link: without an edge these two share a rank, which in a TB
        %% subgraph puts them side by side instead of stacked
        Q1 ~~~ N1
    end
    subgraph C2["Call 2 - read those 3 properly"]
        direction TB
        Q2["<b>Choice:</b> which of the 3?<br/><i>with real detail now</i>"]
        N2["<b>Nouls:</b> does each one<br/>really do it?"]
        Q2 ~~~ N2
    end
    REQ["the request"] --> C1
    C1 -->|"top 3"| C2
    C1 -->|"nothing<br/>applies"| STOP["suggest<br/>nothing"]
    C2 -->|"none fit"| STOP
    C2 -->|"a winner"| OUT["suggest<br/>the winner"]

설정

  • TypeSafe 클라이언트, Anthropic 클라이언트, 그리고 공용 cookbook 헬퍼를 설치합니다.
  • TypeSafe API 키를 설정하고, 측정 대상 에이전트를 위한 Anthropic 키도 설정합니다.
pip install anthropic matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
export TYPESAFE_API_KEY=your-key-here
export ANTHROPIC_API_KEY=your-key-here

참고: 아래 코드 블록은 순서대로 이어지는 하나의 스크립트입니다. 따라 하려면 표시된 순서대로 하나의 파일에 넣으십시오.

결과 캐싱

JsonCache는 각 호출의 결과를 입력을 키로 저장하므로, 다시 실행하면 어느 API도 호출하지 않고 아래 숫자를 재생합니다. 실시간으로 실행하려면 json_cache.json을 삭제하십시오. 공개된 실행은 jev-1.12와 claude-haiku-4-5-20251001을 사용했으며, 2026-07-31에 렌더링되었습니다.

import json
import os
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from time import perf_counter

import anthropic
import matplotlib
import matplotlib.pyplot as plt
from matplotlib.ticker import PercentFormatter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient

matplotlib.use("Agg")  # headless render

TYPESAFE_MODEL = "jev-1.12"
AGENT_MODEL = (
    "claude-haiku-4-5-20251001"  # the agent under test, pinned so scores are stable
)

SHORTLIST = 3  # candidates carried from the first request into the second
EXCERPT_CHARS = (
    700  # SKILL.md characters each candidate brings; the roster file stores 1600
)
GATE_THRESHOLD = (
    0.30  # mean of the three request nouls, below which nothing is suggested
)
FITS_THRESHOLD = (
    0.30  # a shortlist whose best "does this fit" noul is under this is dropped
)
WORKERS = 8  # small pool: enough to keep a live run to minutes, gentle on rate limits

assert EXCERPT_CHARS <= 1600, (
    "the shipped roster file stores 1600 body characters per skill"
)

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
agent = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "cache-only"))
json_cache = JsonCache(Path("json_cache.json"))

1단계: 로스터 로드

hermes_roster.json에는 NousResearch/hermes-agent(MIT)의 고정된 커밋 하나에 있는 182개 스킬이 들어 있습니다. 각 레코드는 스킬의 이름과 분류, 인덱스에 표시되는 설명, 전체 설명, 그리고 해당 SKILL.md의 시작 부분을 담고 있습니다.

아래 인덱스와, 프롬프트에서 그 위에 있는 지시문은 Hermes에서 그대로 복사한 것입니다.

ROSTER = json.loads(Path("hermes_roster.json").read_text(encoding="utf-8"))
BY_NAME = {skill["name"]: skill for skill in ROSTER}

# Verbatim from hermes-agent agent/prompt_builder.py:build_skills_system_prompt.
PREAMBLE = (
    "## Skills (mandatory)\n"
    "Before replying, scan the skills below. If a skill matches or is even partially relevant "
    "to your task, you MUST load it with skill_view(name) and follow its instructions. "
    "Err on the side of loading — it is always better to have context you don't need "
    "than to miss critical steps, pitfalls, or established workflows. "
    "Skills contain specialized knowledge — API endpoints, tool-specific commands, "
    "and proven workflows that outperform general-purpose approaches. Load the skill "
    "even if you think you could handle the task with basic tools like web_search or terminal. "
    "Skills also encode the user's preferred approach, conventions, and quality standards "
    "for tasks like code review, planning, and testing — load them even for tasks you "
    "already know how to do, because the skill defines how it should be done here.\n"
    "Whenever the user asks you to configure, set up, install, enable, disable, modify, "
    "or troubleshoot Hermes Agent itself — its CLI, config, models, providers, tools, "
    "skills, voice, gateway, plugins, or any feature — load the `hermes-agent` skill "
    "first. It has the actual commands (e.g. `hermes config set …`, `hermes tools`, "
    "`hermes setup`) so you don't have to guess or invent workarounds.\n"
    "If a skill has issues, fix it with skill_manage(action='patch').\n"
    "After difficult/iterative tasks, offer to save as a skill. "
    "If a skill you loaded was missing steps, had wrong commands, or needed "
    "pitfalls you discovered, update it before finishing.\n"
    "\n"
)
FOOTER = "\n\nOnly proceed without loading a skill if genuinely none are relevant to the task."
IDENTITY = (
    "You are Hermes, a capable AI assistant with access to tools and a library "
    "of skills. You help the user with coding, research, and everyday tasks.\n\n"
)

def render_index() -> str:
    """The body of <available_skills>: skills grouped by category, both sorted by name."""
    by_category = defaultdict(list)
    for skill in ROSTER:
        by_category[skill["category"]].append(skill)
    lines = []
    for category in sorted(by_category):
        lines.append(f"  {category}:")
        for skill in sorted(by_category[category], key=lambda s: s["name"]):
            lines.append(f"    - {skill['name']}: {skill['description']}")
    return "\n".join(lines)

CATALOG_PROMPT = (
    IDENTITY
    + PREAMBLE
    + "<available_skills>\n"
    + render_index()
    + "\n</available_skills>"
    + FOOTER
)

widths = [len(skill["description"]) for skill in ROSTER]
print(f"{len(ROSTER)} skills in {len({s['category'] for s in ROSTER})} categories")
print(f"roster prompt: {len(CATALOG_PROMPT):,} characters")
print(
    f"index description: {sum(widths) / len(widths):.0f} characters on average, "
    f"{max(widths)} at most"
)
print("\none category, as the agent reads it:")
index_lines = render_index().splitlines()
start = index_lines.index("  apple:")
end = next(
    i
    for i in range(start + 1, len(index_lines))
    if not index_lines[i].startswith("    ")
)
print("\n".join(index_lines[start:end]))
182 skills in 33 categories
roster prompt: 16,089 characters
index description: 54 characters on average, 60 at most

one category, as the agent reads it:
  apple:
    - apple-notes: Manage Apple Notes via memo CLI: create, search, edit.
    - apple-reminders: Apple Reminders via remindctl: add, list, complete.
    - findmy: Track Apple devices/AirTags via FindMy.app on macOS.
    - imessage: Send and receive iMessages/SMS via the imsg CLI on macOS.

2단계: 에이전트 단독 채점

requests.json에는 488개의 단일 턴 요청이 들어 있습니다. 그중 315개는 정확히 하나의 스킬이 담당하고, 나머지 173개는 아무 스킬도 담당하지 않습니다.

담당 스킬이 있는 요청은 Claude Sonnet 5가 각 스킬 자신의 SKILL.md를 바탕으로 작성했으므로 레이블을 신뢰할 수 있고, 요청도 사용자가 실제로 보내는 것보다 쉽습니다.

담당 스킬이 없는 173개는 모두 추측을 벌하기 위해 작성되었습니다. 일상적인 요청 85개, 어떤 스킬도 감당하지 못하는 기술 질문 42개(monad가 무엇인지 설명해 줘), 그리고 로스터에 해당 스킬이 없는 특정한 무언가를 요구하는 46개입니다. 예를 들어 X만 담당하고 다른 건 전혀 없는 로스터에 이걸 Mastodon에 올려 줘 같은 요청입니다.

채점은 에이전트의 첫 응답만 읽습니다. 두 숫자 모두 오류율이므로 각각 낮을수록 좋습니다.

  • 잘못된 로드: 담당 스킬이 있는 요청 중에서, 첫 skill_view 호출이 담당 스킬이 아니었던 비율입니다. 아무것도 로드하지 않은 턴도 실패로 셉니다.
  • 불필요한 로드: 담당 스킬이 없는 요청 중에서, 에이전트가 skill_view를 아예 호출한 비율입니다.
REQUESTS = json.loads(Path("requests.json").read_text(encoding="utf-8"))
POSITIVES = [p for p in REQUESTS if p["gold"]]
NEGATIVES = [p for p in REQUESTS if not p["gold"]]

print(
    f"{len(REQUESTS)} requests: {len(POSITIVES)} covered by a skill "
    f"({len({p['gold'] for p in POSITIVES})} distinct skills), {len(NEGATIVES)} covered by none"
)
print(f"\ncovered   [{POSITIVES[0]['gold']}]  {POSITIVES[0]['text']}")
print(f"uncovered  {NEGATIVES[0]['text']}")
488 requests: 315 covered by a skill (171 distinct skills), 173 covered by none

covered   [1password]  I've got a config.yaml with `{{ op://app-prod/db/password }}` placeholders in it — can you set up my project to pull the real values in at runtime instead of hardcoding them?
uncovered  Add these three cards to our Trello backlog.

제안은 로스터 안이 아니라 그 뒤에, 시스템 프롬프트의 별도 블록으로 들어갑니다. 그래야 매 턴 로스터 텍스트가 동일하게 유지되어 프리픽스 캐싱이 보존됩니다.

에이전트는 최소한의 도구 집합을 가지는데, 여기에는 자유 텍스트 이름으로 스킬을 로드하는 skill_view가 포함됩니다. 올바른 로드가 되려면 이름이 스킬과 정확히 일치해야 합니다.

# Verbatim from hermes-agent tools/skills_tool.py:SKILL_VIEW_SCHEMA.
SKILL_VIEW_DESCRIPTION = (
    "Skills allow for loading information about specific tasks and workflows, as "
    "well as scripts and templates. Load a skill's full content or access its "
    "linked files (references, templates, scripts). First call returns SKILL.md "
    "content plus a 'linked_files' dict showing available references/templates/"
    "scripts. To access those, call again with file_path parameter."
)
TOOLS = [
    {
        "name": "skill_view",
        "description": SKILL_VIEW_DESCRIPTION,
        "input_schema": {
            "type": "object",
            "properties": {
                "name": {"type": "string", "description": "The skill name."}
            },
            "required": ["name"],
        },
    },
    {
        "name": "terminal",
        "description": "Run a shell command on the user's machine and return its output.",
        "input_schema": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
    {
        "name": "read_file",
        "description": "Read a file from the user's filesystem.",
        "input_schema": {
            "type": "object",
            "properties": {"path": {"type": "string"}},
            "required": ["path"],
        },
    },
    {
        "name": "web_search",
        "description": "Search the web and return result snippets.",
        "input_schema": {
            "type": "object",
            "properties": {"query": {"type": "string"}},
            "required": ["query"],
        },
    },
]

@json_cache
def run_turn(model: str, arm: str, request: str, suggestion: str) -> dict:
    """One measured turn. ``arm`` is in the key so each arm samples independently."""
    system = [
        {"type": "text", "text": CATALOG_PROMPT, "cache_control": {"type": "ephemeral"}}
    ]
    if suggestion:
        system.append({"type": "text", "text": suggestion})  # after the breakpoint
    response = agent.messages.create(
        model=model,
        max_tokens=1024,
        system=system,
        tools=TOOLS,
        messages=[{"role": "user", "content": request}],
    )
    usage = response.usage
    return {
        "loaded": [
            str(block.input.get("name", ""))
            for block in response.content
            if block.type == "tool_use" and block.name == "skill_view"
        ],
        "input_tokens": usage.input_tokens or 0,
        "output_tokens": usage.output_tokens or 0,
    }

def summarise(turns: dict[str, dict]) -> dict[str, float]:
    """Two failure rates: wrong loads on covered requests, needless ones on uncovered."""
    hits = [turns[p["text"]]["loaded"][:1] == [p["gold"]] for p in POSITIVES]
    over = [bool(turns[p["text"]]["loaded"]) for p in NEGATIVES]
    return {
        # both metrics are errors, so the two columns read the same direction
        "wrong_load": 1 - sum(hits) / len(hits),
        "needless_load": sum(over) / len(over),
    }

def run_arm(arm: str, suggestions: dict[str, str]) -> dict[str, dict]:
    """One measured turn per request, in a small pool. 488 calls."""
    texts = [request["text"] for request in REQUESTS]
    with ThreadPoolExecutor(max_workers=WORKERS) as pool:
        turns = pool.map(
            lambda t: run_turn(AGENT_MODEL, arm, t, suggestions.get(t, "")), texts
        )
        return dict(zip(texts, turns))

에이전트는 오늘날 작동하는 방식대로, 로스터만 가지고 먼저 실행됩니다. 이때의 두 오류율이 cookbook 나머지가 비교 기준으로 삼는 베이스라인입니다.

baseline = run_arm("baseline", {})
base_scores = summarise(baseline)
print(
    f"wrong loads    {base_scores['wrong_load']:.1%}   ({len(POSITIVES)} covered requests)"
)
print(
    f"needless loads {base_scores['needless_load']:.1%}   ({len(NEGATIVES)} uncovered requests)"
)

# where the wrong loads land: a neighbour of the right skill, or somewhere unrelated?
misses = [
    (p["gold"], baseline[p["text"]]["loaded"][0])
    for p in POSITIVES
    if baseline[p["text"]]["loaded"] and baseline[p["text"]]["loaded"][0] != p["gold"]
]
same_category = sum(
    1
    for gold, got in misses
    if got in BY_NAME and BY_NAME[got]["category"] == BY_NAME[gold]["category"]
)
print(
    f"\nof {len(misses)} wrong first picks, {same_category} came from the right skill's own "
    f"category"
)
wrong loads    16.8%   (315 covered requests)
needless loads 9.8%   (173 uncovered requests)

of 36 wrong first picks, 10 came from the right skill's own category

잘못된 로드는 우연히 그럴 확률보다 훨씬 자주 올바른 스킬 자신의 분류에 떨어집니다. 따라서 어려운 부분은 몇몇 닮은꼴을 구별하는 것입니다. 에이전트는 이미 대략 올바른 위치를 살펴보고 있습니다.

3단계: 전체 로스터 순위 매기기

하나의 요청은 두 종류의 질문을 담습니다.

  • **which**는 182개 스킬 이름 전체를 대상으로 하는 Choice 질문이며, 인덱스 설명이 각 선택지의 criteria입니다 (에이전트 자신이 받는 것과 같은 텍스트입니다). 그 확률이 곧 순위입니다.
  • 요청에 관한 세 개의 Noul 질문은 아래에 나열되며, 각각 다른 방식으로 설명이 아니라 행동을 원하는지 묻습니다. prose_suffices는 방향이 반대입니다. 이들의 평균이 무언가를 제안할지 전혀 아닌지를 결정하며, 0.30 미만이면 아무것도 제안하지 않습니다.

둘 모두 하나의 요청으로 나가므로, 순위 매기기와 검사에 왕복 한 번만 듭니다.

이 세 질문은 행동이 필요한지를 묻도록 작성하십시오. 주제 내용에 관한 질문은 monad가 무엇인지 설명해 줘를 스킬이 필요한 요청과 구분하지 못합니다. 둘 다 소프트웨어이기 때문입니다.

Choice 질문 하나로 이 정도 크기의 로스터는 여유롭게 담깁니다. 몇 배만 더 커지면 그것을 여러 청크로 나눠 각각 순위를 매기고, 그다음 이와 같은 숏리스트 단계를 승자들에게 실행하게 될 것입니다.

CHOICE_INSTRUCTIONS = (
    "Which of these skills, if any, is the right one to load to help with the "
    "user's latest request?"
)
GATE_QUESTIONS = {
    "acts_on_user_system": (
        "Is the assistant being asked to act on the user's files, accounts, devices, "
        "or online services, rather than only to explain or advise?"
    ),
    "would_follow_documented_procedure": (
        "Would a careful expert answering this consult a specific documented procedure "
        "or set of commands, rather than answering from general understanding?"
    ),
    "prose_suffices": (
        "Could a knowledgeable generalist fully satisfy this request in prose, with "
        "no tools, no documentation, and no access to the user's files or accounts?"
    ),
}
INVERTED = {"prose_suffices"}  # a yes here points away from needing a skill

def build_state(request: str) -> dict:
    return {"request": request, "recent_context": ""}

@json_cache
def rank_wide(request: str) -> dict:
    """Request 1: rank all 182 skills, and score the request for whether a skill applies."""
    questions = {
        "which": Choice(
            instructions=CHOICE_INSTRUCTIONS,
            criteria={skill["name"]: skill["description"] for skill in ROSTER},
        )
    }
    for key, text in GATE_QUESTIONS.items():
        questions[f"gate::{key}"] = Noul(instructions=text)
    started = perf_counter()
    response = client.system_one(
        state=build_state(request), questions=questions, model=TYPESAFE_MODEL
    )
    ranked = sorted(
        response.answers["which"].probabilities.items(), key=lambda kv: -kv[1]
    )
    values = {
        key.removeprefix("gate::"): answer.noul
        for key, answer in response.answers.items()
        if key.startswith("gate::")
    }
    oriented = [(1.0 - v) if k in INVERTED else v for k, v in values.items()]
    return {
        "ranked": ranked[
            :12
        ],  # more than any shortlist needs, and keeps the cache small
        "gate": sum(oriented) / len(oriented),
        "values": values,
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

DEMO = [
    "Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so it syncs"
    " to my phone? Just write it up in whatever editor pops up.",
    "Can you put together a pitch deck skeleton (cover, situation overview, comps, precedent"
    " transactions, DCF, LBO) as a .pptx, using our firm-template.pptx for branding and"
    " footnoting each valuation number back to the cell it came from in the model?",
    "Post this announcement to my Mastodon account.",
]
for request in DEMO:
    wide = rank_wide(request)
    verdict = "suggest" if wide["gate"] >= GATE_THRESHOLD else "stay quiet"
    print(f'"{request[:78]}"')
    print(f"  needs a skill {wide['gate']:.2f} -> {verdict}   ({wide['seconds']}s)")
    for name, probability in wide["ranked"][:SHORTLIST]:
        print(f"    {probability:.3f}  {name:<38}{BY_NAME[name]['description']}")
    print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
  needs a skill 0.75 -> suggest   (0.31s)
    0.990  apple-notes                           Manage Apple Notes via memo CLI: create, search, edit.
    0.010  computer-use                          Drive the user's desktop in the background — clicking, ty...
    0.000  concept-diagrams                      Generate flat, minimal educational SVG visuals as HTML.

"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
  needs a skill 0.76 -> suggest   (0.16s)
    0.700  powerpoint                            Create, read, edit .pptx decks, slides, notes, templates.
    0.300  pptx-author                           Build PowerPoint decks headless with python-pptx.
    0.000  chroma                                Embedding database for RAG and semantic search.

"Post this announcement to my Mastodon account."
  needs a skill 0.78 -> suggest   (0.16s)
    0.550  xurl                                  X/Twitter via xurl CLI: raw post search, posting, DM, media.
    0.140  computer-use                          Drive the user's desktop in the background — clicking, ty...
    0.080  openhands                             Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).

Notes.app 요청은 명확하고 1위 선택지가 정답입니다. Mastodon 쪽은 순위 매기기로 구할 수 없습니다. 세 질문 모두 스킬이 필요하다고 말합니다. 계정에 게시하는 것은 행동이고, X에 게시하는 스킬은 있지만 Mastodon용은 없으니 가장 가까운 스킬이 어쨌든 이깁니다.

남은 것은 덱입니다. 선두 둘 모두 .pptx 스킬이고, 60자에서는 이 광범위한 Choice 질문이 편집 스킬을 작성 스킬보다 앞에 놓습니다. 덱을 작성하는 요청인데도 그렇습니다.

4단계: 상위 세 개 재정렬

세 개의 선택지는 전체 설명에 각 스킬 자신의 SKILL.md 시작 부분까지 넣을 여유를 줍니다. 그래서 두 번째 요청은 같은 질문을 더 나은 증거 위에서 다시 던집니다.

  • **which**는 숏리스트를 대상으로 하는 Choice 질문이며, 그 더 긴 텍스트가 각 선택지의 criteria입니다.
  • **fits::{name}**은 후보마다 하나씩 있는 Noul 질문입니다. 이 스킬이 요청이 요구하는 특정한 일을 하는가? 각각 독립적으로 답하므로 전부 낮게 나올 수 있으며, 가장 높은 값이 0.30 미만인 숏리스트는 통째로 버려집니다.
RERANK_INSTRUCTIONS = (
    "Exactly one of these skills is the right one to load for the user's latest "
    "request. Which one? Read what each actually does, not just its name."
)

def rerank_criteria(names: tuple[str, ...], excerpt: int) -> dict[str, str]:
    return {
        name: f"{BY_NAME[name]['description_full']} — {BY_NAME[name]['body'][:excerpt]}"
        for name in names
    }

def rerank_questions(names: tuple[str, ...], excerpt: int) -> dict:
    questions = {
        "which": Choice(
            instructions=RERANK_INSTRUCTIONS, criteria=rerank_criteria(names, excerpt)
        )
    }
    for name in names:
        questions[f"fits::{name}"] = Noul(
            instructions=(
                f"Does the skill '{name}' do the specific thing the user's request asks "
                f"for? It is described as: {BY_NAME[name]['description_full']}"
            )
        )
    return questions

@json_cache
def rerank(request: str, names: tuple[str, ...], excerpt: int) -> dict:
    """Request 2: the same Choice over a shortlist, plus one absolute noul per candidate."""
    started = perf_counter()
    response = client.system_one(
        state=build_state(request),
        questions=rerank_questions(names, excerpt),
        model=TYPESAFE_MODEL,
    )
    return {
        "winner": response.answers["which"].choice,
        "fits": {
            key.removeprefix("fits::"): answer.noul
            for key, answer in response.answers.items()
            if key.startswith("fits::")
        },
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

for request in DEMO:
    wide = rank_wide(request)
    if wide["gate"] < GATE_THRESHOLD:
        print(f'"{request[:78]}"\n  scored too low, nothing suggested\n')
        continue
    shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
    result = rerank(request, shortlist, EXCERPT_CHARS)
    best = max(result["fits"].values())
    verdict = result["winner"] if best >= FITS_THRESHOLD else "nothing fits"
    print(f'"{request[:78]}"')
    print(f"  was {shortlist[0]} -> {verdict}   ({result['seconds']}s)")
    for name in shortlist:
        print(f"    fits {result['fits'][name]:.2f}  {name}")
    print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
  was apple-notes -> apple-notes   (0.12s)
    fits 0.60  apple-notes
    fits 0.54  computer-use
    fits 0.01  concept-diagrams

"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
  was powerpoint -> pptx-author   (0.09s)
    fits 0.73  powerpoint
    fits 0.38  pptx-author
    fits 0.02  chroma

"Post this announcement to my Mastodon account."
  was xurl -> xurl   (0.09s)
    fits 0.56  xurl
    fits 0.38  computer-use
    fits 0.05  openhands

두 .pptx 스킬은 각자 자신의 텍스트를 가져오자 구분됩니다. 덱 요청은 작성 스킬로 뒤집힙니다.

여기서 fits noul과 Choice는 의견이 갈립니다. noul은 편집 스킬에 더 높은 점수를 주는데 Choice는 작성 스킬을 고릅니다. 둘은 서로 다른 것을 결정하고 있습니다. Choice는 어느 스킬인지를, noul은 아예 말할지 여부를 정합니다.

Mastodon 요청은 두 검사를 모두 통과합니다. 가장 좋은 fits noul이 0.30을 넘으므로, 이 레시피는 Mastodon에 관한 요청에 X 스킬을 제안합니다. 이와 비슷한 대부분의 요청은 걸러집니다. 두 번째 단계는 광범위한 순위 매기기가 넘겨준 것만 거부할 수 있는데, 여기서는 세 개의 아슬아슬한 후보였습니다.

아래 함수가 레시피 전체입니다. 두 번의 요청과 두 개의 임계값, 그리고 최대 하나의 스킬 이름이 돌아옵니다.

자신의 로스터를 향하게 하려면 hermes_roster.json을 교체하십시오. 위의 모든 질문은 그 파일에서 name, description, description_full, body만 읽으며, 그 밖에 Hermes를 아는 곳은 없습니다.

def suggest(request: str) -> tuple[str, ...]:
    """At most one skill name for a request, or () for "nothing here applies"."""
    wide = rank_wide(request)
    if wide["gate"] < GATE_THRESHOLD:
        return ()
    shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
    result = rerank(request, shortlist, EXCERPT_CHARS)
    if max(result["fits"].values()) < FITS_THRESHOLD:
        return ()
    return (result["winner"],)

def suggestion_block(names: tuple[str, ...]) -> str:
    """What gets appended after the roster, in the suggestion.

    This string is a measured input rather than prose: it goes to the agent, so it is part
    of every graded turn's cache key. Editing a word here silently invalidates the shipped
    results and costs a live re-run to restore them.
    """
    body = (
        f"Relevant to the current request: {', '.join(names)}. Ignore this if it does not "
        "fit what the user actually asked for."
        if names
        else "No skill in the roster appears relevant to this request."
    )
    return f"\n\n<skill_relevance>\n{body}\n</skill_relevance>"

print(suggestion_block(suggest(DEMO[1])))
print(suggestion_block(suggest(DEMO[2])))

<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>

<skill_relevance>
Relevant to the current request: xurl. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>

5단계: 제안 측정

488개 요청 각각이 에이전트에 세 번 전달되며, 매번 하나의 측정 턴입니다. 실행들은 에이전트에게 무엇을 알려주는지만 다릅니다.

시스템 프롬프트에 들어가는 것
단독 에이전트 아무것도 없음
제안이 있는 에이전트 suggest()가 반환한 것
정답을 받은 에이전트 담당 스킬의 이름, 없을 때는 “적용되는 것이 없음”

세 번째는 달성할 수 없습니다. 다른 둘을 견줘 측정하는 천장입니다.

그 제안의 문구는 두 가지 일을 합니다. 제안은 무시해도 된다고 말합니다. 더 세게 밀어붙이면 잘못된 제안에도 순응을 얻어내기 때문이며, 잘못된 제안은 아예 없는 것보다 나쁩니다. 그리고 제안할 것이 없는 턴에서도 그렇다고 말하는 문장을 보냅니다. 아무것도 보내지 않으면 로스터 자체의 “로드하는 쪽으로 기울라”는 지시를 견제할 것이 없어집니다.

texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:  # up to 488 x 2 TypeSafe requests
    suggested = dict(zip(texts, pool.map(suggest, texts)))
WIDE = {text: rank_wide(text) for text in texts}  # all cache hits now; reused below

arms = {
    "baseline": {},
    "TypeSafe": {
        request["text"]: suggestion_block(suggested[request["text"]])
        for request in REQUESTS
    },
    "oracle": {
        request["text"]: suggestion_block((request["gold"],) if request["gold"] else ())
        for request in REQUESTS
    },
}
scores = {
    arm: summarise(run_arm(arm, suggestions)) for arm, suggestions in arms.items()
}

print(f"{'run':<10}{'wrong loads':>13}{'needless loads':>16}")
for arm, row in scores.items():
    print(f"{arm:<10}{row['wrong_load']:>13.1%}{row['needless_load']:>16.1%}")

def fewer(metric: str) -> str:
    """The plain ratio between the two arms' error rates."""
    return f"{scores['baseline'][metric] / scores['TypeSafe'][metric]:.1f}x fewer"

print(
    f"\nbaseline -> TypeSafe:  {fewer('wrong_load')} wrong loads, "
    f"{fewer('needless_load')} needless ones"
)
run         wrong loads  needless loads
baseline          16.8%            9.8%
TypeSafe           7.3%            4.0%
oracle             2.5%            1.2%

baseline -> TypeSafe:  2.3x fewer wrong loads, 2.4x fewer needless ones
moved = [
    (
        baseline[p["text"]]["loaded"][:1] == [p["gold"]],
        run_turn(AGENT_MODEL, "TypeSafe", p["text"], arms["TypeSafe"][p["text"]])[
            "loaded"
        ][:1]
        == [p["gold"]],
    )
    for p in POSITIVES
]
print(
    f"of {len(POSITIVES)} covered requests: {sum(not b and a for b, a in moved)} the suggestion "
    f"fixed, {sum(b and not a for b, a in moved)} it broke"
)
of 315 covered requests: 37 the suggestion fixed, 7 it broke

제안은 망가뜨리는 요청보다 고치는 요청이 훨씬 많지만, 에이전트가 혼자서 맞혔던 일부를 실제로 망가뜨립니다. 확신에 찬 잘못된 제안은 아예 없는 제안보다 더 설득력이 큽니다. 이것이 제안을 턴 앞에 놓는 대가입니다.

SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"

ARM_COLOR = {"baseline": BLUE, "TypeSafe": ORANGE, "oracle": MUTED}

def style(ax):
    ax.set_facecolor(SURFACE)
    for side in ("top", "right"):
        ax.spines[side].set_visible(False)
    for side in ("left", "bottom"):
        ax.spines[side].set_color(AXIS)
    ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
    ax.set_axisbelow(True)

panels = [
    ("wrong_load", f"wrong loads\n{len(POSITIVES)} covered requests"),
    ("needless_load", f"needless loads\n{len(NEGATIVES)} uncovered requests"),
]
names = list(scores)
fig, axes = plt.subplots(1, 2, figsize=(8.4, 3.6), facecolor=SURFACE)
for ax, (metric, title) in zip(axes, panels):
    style(ax)
    ax.grid(axis="y", color=GRID, linewidth=0.8)
    values = [scores[arm][metric] for arm in names]
    bars = ax.bar(
        names,
        values,
        0.58,
        color=[ARM_COLOR[arm] for arm in names],
        # the oracle is a ceiling, not a competitor: gray, and hatched so it never depends
        # on colour alone
        hatch=["", "", "///"],
        edgecolor=SURFACE,
        linewidth=1.2,
    )
    ax.bar_label(
        bars,
        labels=[f"{v:.1%}" for v in values],
        padding=3,
        color=INK2,
        fontsize=9,
    )
    ax.set_title(title, loc="left", color=INK2, fontsize=9.5)
    ax.set_ylim(0, max(values) * 1.28)
    ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))
    ax.set_ylabel("% of those requests - lower is better", color=INK2, fontsize=9)
fig.suptitle(
    f"Hermes' {len(ROSTER)}-skill roster, {len(REQUESTS)} requests, {AGENT_MODEL}",
    x=0.02,
    ha="left",
    color=INK,
    fontsize=11,
)
fig.tight_layout()
display(fig)
plt.close(fig)
출력

결과가 보여주는 것

  • 잘못된 로드는 16.8%에서 7.3%로, 불필요한 로드는 9.8%에서 4.0%로 떨어졌습니다. 이는 잘린 인덱스로 추측하는 것과 정답을 받는 것 사이의 격차 대부분을 메웁니다.
  • 에이전트가 혼자서 맞혔던 일부 요청은 제안이 붙으면 오답으로 돌아옵니다. 개수는 위에 있습니다.

자신의 에이전트가 큰 로스터를 안고 있을 때 이 형태를 그대로 쓰십시오. 모든 것에 대한 저렴한 순위 매기기, 그다음 두세 개를 자세히 들여다보기입니다. 어느 단계든 빈손으로 돌아올 수 있습니다.

Playground에서 열기

4단계의 덱 요청에 대한 Playground 링크를 만드십시오. 각 후보의 전체 설명과 본문 발췌를 그 criteria로 사용합니다.

demo_shortlist = tuple(name for name, _ in rank_wide(DEMO[1])["ranked"][:SHORTLIST])
playground_link = make_playground_link(
    build_state(DEMO[1]),
    rerank_questions(demo_shortlist, EXCERPT_CHARS),
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open the shortlist + questions in the TypeSafe playground]({playground_link})"
    )
)
TypeSafe Playground에서 숏리스트 + 질문 열기 →

다음 단계

같은 형태가 다른 곳에서도 나타납니다. 의도 라우팅은 스킬이 아니라 핸들러로 라우팅하는 방법이고, 신뢰도는 두 임계값을 고르는 방법이며, 투기적 팬아웃은 모든 질문을 하나의 요청에 넣는 방법입니다.