Документация

Подсказка навыка

Выбирает не более одного навыка на ход агента из 182 в каталоге Hermes от Nous Research: два запроса TypeSafe ранжируют кандидатов и перепроверяют лучших.

Агенты выбирают навыки, обрезая их и загружая все в системное сообщение, что увеличивает расходы, ухудшает качество выбора навыка и вызывает context rot для остальной части сессии. Мы решаем это двумя запросами TypeSafe на ход — один ранжирует навыки, другой проверяет выбор, — и сокращаем неверные загрузки навыков более чем вдвое.

У агента с большим списком навыков выбор делается почти без информации. Список доходит до него как индекс: одна строка на навык, с обрезанным описанием, чтобы полный текст не вытеснял разговор. Hermes, используемый здесь agent harness, по умолчанию режет его до 60 символов. Например, при такой ширине навык, который редактирует файлы .pptx, читается почти так же, как тот, который их создаёт. Попросите питч-дек — и агент может загрузить не тот. А на ходе, где не подходит ни один навык, он может всё равно загрузить один, потому что список имён провоцирует угадывание.

Этот cookbook не трогает описания и вместо этого применяет progressive disclosure: дёшево читает все 182 навыка, а затем подробно читает три из них. Перед решением о том, какой навык загружать (если вообще загружать), ставятся два запроса TypeSafe. Первый ранжирует каждый навык из списка относительно хода пользователя и отвечает, нужен ли ходу навык вообще. Второй перечитывает только три лучших — теперь с полным описанием каждого навыка и началом его инструкций — и волен отклонить их все.

Имя победителя попадает одной дополнительной строкой в системный промпт агента на этот ход:

<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user
actually asked for.
</skill_relevance>

Агент сохраняет свой полный индекс и собственное суждение, а эта одна строка лишь подсказывает, на какую запись посмотреть первой. Сам список никогда не меняется, поэтому любое префиксное кэширование по нему продолжает работать. На 488 запросах к claude-haiku-4-5-20251001 с навыками из списка Hermes:

загружает не тот навык загружает навык, когда ничего не подходит
агент один, только со своим списком 16.8% 9.8%
агент с подсказкой TypeSafe 7.3% 4.0%
агент, которому дали правильный ответ 2.5% 1.2%

Третья строка показывает, что нижняя граница ошибок не равна нулю: даже получив правильный навык, агент не всегда его загружает, и никакой метод выбора, каким бы хорошим он ни был, этого не преодолевает.

В итоге у вас есть функция suggest(), возвращающая не более одного имени навыка, suggestion_block(), оборачивающая его для системного промпта, и harness, который выдал таблицу выше и готов работать с вашим собственным списком.

flowchart LR
    subgraph C1["Call 1 - skim all 182 skills"]
        direction TB
        Q1["<b>Choice:</b> which skill fits?<br/><i>all 182, one line each</i>"]
        N1["<b>Nouls:</b> need a skill at all?<br/>· act on their stuff?<br/>· follow written steps?<br/>· or just talk?"]
        %% invisible link: without an edge these two share a rank, which in a TB
        %% subgraph puts them side by side instead of stacked
        Q1 ~~~ N1
    end
    subgraph C2["Call 2 - read those 3 properly"]
        direction TB
        Q2["<b>Choice:</b> which of the 3?<br/><i>with real detail now</i>"]
        N2["<b>Nouls:</b> does each one<br/>really do it?"]
        Q2 ~~~ N2
    end
    REQ["the request"] --> C1
    C1 -->|"top 3"| C2
    C1 -->|"nothing<br/>applies"| STOP["suggest<br/>nothing"]
    C2 -->|"none fit"| STOP
    C2 -->|"a winner"| OUT["suggest<br/>the winner"]

Настройка

  • Установите клиент TypeSafe, клиент Anthropic и общие вспомогательные модули cookbook.
  • Задайте ключ API TypeSafe и ключ Anthropic для измеряемого агента.
pip install anthropic matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
export TYPESAFE_API_KEY=your-key-here
export ANTHROPIC_API_KEY=your-key-here

Примечание. Приведённые ниже блоки кода — это один скрипт, по порядку. Чтобы повторить, поместите их в один файл в показанном порядке.

Кэширование результатов

JsonCache сохраняет результат каждого вызова с ключом по его входным данным, поэтому повторный запуск воспроизводит числа ниже, а не обращается к тому или иному API. Удалите json_cache.json, чтобы запустить вживую. Опубликованный запуск использовал jev-1.12 и claude-haiku-4-5-20251001, отрендерено 2026-07-31.

import json
import os
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from time import perf_counter

import anthropic
import matplotlib
import matplotlib.pyplot as plt
from matplotlib.ticker import PercentFormatter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient

matplotlib.use("Agg")  # headless render

TYPESAFE_MODEL = "jev-1.12"
AGENT_MODEL = (
    "claude-haiku-4-5-20251001"  # the agent under test, pinned so scores are stable
)

SHORTLIST = 3  # candidates carried from the first request into the second
EXCERPT_CHARS = (
    700  # SKILL.md characters each candidate brings; the roster file stores 1600
)
GATE_THRESHOLD = (
    0.30  # mean of the three request nouls, below which nothing is suggested
)
FITS_THRESHOLD = (
    0.30  # a shortlist whose best "does this fit" noul is under this is dropped
)
WORKERS = 8  # small pool: enough to keep a live run to minutes, gentle on rate limits

assert EXCERPT_CHARS <= 1600, (
    "the shipped roster file stores 1600 body characters per skill"
)

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
agent = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "cache-only"))
json_cache = JsonCache(Path("json_cache.json"))

Шаг 1: загрузите список

hermes_roster.json содержит 182 навыка из NousResearch/hermes-agent (MIT) на одном закреплённом коммите. Каждая запись хранит имя и категорию навыка, описание в том виде, как его показывает индекс, полное описание и начало его SKILL.md.

Индекс ниже и предшествующие ему в промпте инструкции скопированы из Hermes.

ROSTER = json.loads(Path("hermes_roster.json").read_text(encoding="utf-8"))
BY_NAME = {skill["name"]: skill for skill in ROSTER}

# Verbatim from hermes-agent agent/prompt_builder.py:build_skills_system_prompt.
PREAMBLE = (
    "## Skills (mandatory)\n"
    "Before replying, scan the skills below. If a skill matches or is even partially relevant "
    "to your task, you MUST load it with skill_view(name) and follow its instructions. "
    "Err on the side of loading — it is always better to have context you don't need "
    "than to miss critical steps, pitfalls, or established workflows. "
    "Skills contain specialized knowledge — API endpoints, tool-specific commands, "
    "and proven workflows that outperform general-purpose approaches. Load the skill "
    "even if you think you could handle the task with basic tools like web_search or terminal. "
    "Skills also encode the user's preferred approach, conventions, and quality standards "
    "for tasks like code review, planning, and testing — load them even for tasks you "
    "already know how to do, because the skill defines how it should be done here.\n"
    "Whenever the user asks you to configure, set up, install, enable, disable, modify, "
    "or troubleshoot Hermes Agent itself — its CLI, config, models, providers, tools, "
    "skills, voice, gateway, plugins, or any feature — load the `hermes-agent` skill "
    "first. It has the actual commands (e.g. `hermes config set …`, `hermes tools`, "
    "`hermes setup`) so you don't have to guess or invent workarounds.\n"
    "If a skill has issues, fix it with skill_manage(action='patch').\n"
    "After difficult/iterative tasks, offer to save as a skill. "
    "If a skill you loaded was missing steps, had wrong commands, or needed "
    "pitfalls you discovered, update it before finishing.\n"
    "\n"
)
FOOTER = "\n\nOnly proceed without loading a skill if genuinely none are relevant to the task."
IDENTITY = (
    "You are Hermes, a capable AI assistant with access to tools and a library "
    "of skills. You help the user with coding, research, and everyday tasks.\n\n"
)

def render_index() -> str:
    """The body of <available_skills>: skills grouped by category, both sorted by name."""
    by_category = defaultdict(list)
    for skill in ROSTER:
        by_category[skill["category"]].append(skill)
    lines = []
    for category in sorted(by_category):
        lines.append(f"  {category}:")
        for skill in sorted(by_category[category], key=lambda s: s["name"]):
            lines.append(f"    - {skill['name']}: {skill['description']}")
    return "\n".join(lines)

CATALOG_PROMPT = (
    IDENTITY
    + PREAMBLE
    + "<available_skills>\n"
    + render_index()
    + "\n</available_skills>"
    + FOOTER
)

widths = [len(skill["description"]) for skill in ROSTER]
print(f"{len(ROSTER)} skills in {len({s['category'] for s in ROSTER})} categories")
print(f"roster prompt: {len(CATALOG_PROMPT):,} characters")
print(
    f"index description: {sum(widths) / len(widths):.0f} characters on average, "
    f"{max(widths)} at most"
)
print("\none category, as the agent reads it:")
index_lines = render_index().splitlines()
start = index_lines.index("  apple:")
end = next(
    i
    for i in range(start + 1, len(index_lines))
    if not index_lines[i].startswith("    ")
)
print("\n".join(index_lines[start:end]))
182 skills in 33 categories
roster prompt: 16,089 characters
index description: 54 characters on average, 60 at most

one category, as the agent reads it:
  apple:
    - apple-notes: Manage Apple Notes via memo CLI: create, search, edit.
    - apple-reminders: Apple Reminders via remindctl: add, list, complete.
    - findmy: Track Apple devices/AirTags via FindMy.app on macOS.
    - imessage: Send and receive iMessages/SMS via the imsg CLI on macOS.

Шаг 2: оцените агента в одиночку

requests.json содержит 488 одноходовых запросов: 315 из них покрываются ровно одним навыком, а остальные 173 — ничем.

Покрытые запросы написал Claude Sonnet 5 по собственному SKILL.md каждого навыка, поэтому метки заслуживают доверия, а сами запросы проще тех, что присылают пользователи.

Все 173 непокрытых запроса написаны, чтобы наказывать угадывание: 85 бытовых запросов, 42 технических вопроса, которые не обслуживает ни один навык (объясни, что такое монада), и 46 запросов о чём-то конкретном, для чего в списке нет навыка, — например, опубликуй это в Mastodon при списке, покрывающем X и ничего больше.

Оценка читает только первый ответ агента. Оба числа — доли ошибок, поэтому меньше — лучше для каждого:

  • Неверная загрузка: из покрытых запросов — доля, где первый вызов skill_view был не тем навыком, который покрывает запрос. Ход, который не загрузил ничего, считается промахом.
  • Лишняя загрузка: из непокрытых запросов — доля, где агент вообще вызвал skill_view.
REQUESTS = json.loads(Path("requests.json").read_text(encoding="utf-8"))
POSITIVES = [p for p in REQUESTS if p["gold"]]
NEGATIVES = [p for p in REQUESTS if not p["gold"]]

print(
    f"{len(REQUESTS)} requests: {len(POSITIVES)} covered by a skill "
    f"({len({p['gold'] for p in POSITIVES})} distinct skills), {len(NEGATIVES)} covered by none"
)
print(f"\ncovered   [{POSITIVES[0]['gold']}]  {POSITIVES[0]['text']}")
print(f"uncovered  {NEGATIVES[0]['text']}")
488 requests: 315 covered by a skill (171 distinct skills), 173 covered by none

covered   [1password]  I've got a config.yaml with `{{ op://app-prod/db/password }}` placeholders in it — can you set up my project to pull the real values in at runtime instead of hardcoding them?
uncovered  Add these three cards to our Trello backlog.

Подсказка идёт отдельным блоком системного промпта — после списка, а не внутри него, — поэтому текст списка одинаков на каждом ходе, что поддерживает префиксное кэширование.

У агента минимальный набор инструментов, включая skill_view для загрузки навыка по имени в свободной форме. Для корректной загрузки имя должно точно совпадать с навыком.

# Verbatim from hermes-agent tools/skills_tool.py:SKILL_VIEW_SCHEMA.
SKILL_VIEW_DESCRIPTION = (
    "Skills allow for loading information about specific tasks and workflows, as "
    "well as scripts and templates. Load a skill's full content or access its "
    "linked files (references, templates, scripts). First call returns SKILL.md "
    "content plus a 'linked_files' dict showing available references/templates/"
    "scripts. To access those, call again with file_path parameter."
)
TOOLS = [
    {
        "name": "skill_view",
        "description": SKILL_VIEW_DESCRIPTION,
        "input_schema": {
            "type": "object",
            "properties": {
                "name": {"type": "string", "description": "The skill name."}
            },
            "required": ["name"],
        },
    },
    {
        "name": "terminal",
        "description": "Run a shell command on the user's machine and return its output.",
        "input_schema": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
    {
        "name": "read_file",
        "description": "Read a file from the user's filesystem.",
        "input_schema": {
            "type": "object",
            "properties": {"path": {"type": "string"}},
            "required": ["path"],
        },
    },
    {
        "name": "web_search",
        "description": "Search the web and return result snippets.",
        "input_schema": {
            "type": "object",
            "properties": {"query": {"type": "string"}},
            "required": ["query"],
        },
    },
]

@json_cache
def run_turn(model: str, arm: str, request: str, suggestion: str) -> dict:
    """One measured turn. ``arm`` is in the key so each arm samples independently."""
    system = [
        {"type": "text", "text": CATALOG_PROMPT, "cache_control": {"type": "ephemeral"}}
    ]
    if suggestion:
        system.append({"type": "text", "text": suggestion})  # after the breakpoint
    response = agent.messages.create(
        model=model,
        max_tokens=1024,
        system=system,
        tools=TOOLS,
        messages=[{"role": "user", "content": request}],
    )
    usage = response.usage
    return {
        "loaded": [
            str(block.input.get("name", ""))
            for block in response.content
            if block.type == "tool_use" and block.name == "skill_view"
        ],
        "input_tokens": usage.input_tokens or 0,
        "output_tokens": usage.output_tokens or 0,
    }

def summarise(turns: dict[str, dict]) -> dict[str, float]:
    """Two failure rates: wrong loads on covered requests, needless ones on uncovered."""
    hits = [turns[p["text"]]["loaded"][:1] == [p["gold"]] for p in POSITIVES]
    over = [bool(turns[p["text"]]["loaded"]) for p in NEGATIVES]
    return {
        # both metrics are errors, so the two columns read the same direction
        "wrong_load": 1 - sum(hits) / len(hits),
        "needless_load": sum(over) / len(over),
    }

def run_arm(arm: str, suggestions: dict[str, str]) -> dict[str, dict]:
    """One measured turn per request, in a small pool. 488 calls."""
    texts = [request["text"] for request in REQUESTS]
    with ThreadPoolExecutor(max_workers=WORKERS) as pool:
        turns = pool.map(
            lambda t: run_turn(AGENT_MODEL, arm, t, suggestions.get(t, "")), texts
        )
        return dict(zip(texts, turns))

Сначала агент работает только со своим списком — так, как он работает сегодня. Две его доли ошибок — это базовый уровень, с которым сравнивается остальной cookbook.

baseline = run_arm("baseline", {})
base_scores = summarise(baseline)
print(
    f"wrong loads    {base_scores['wrong_load']:.1%}   ({len(POSITIVES)} covered requests)"
)
print(
    f"needless loads {base_scores['needless_load']:.1%}   ({len(NEGATIVES)} uncovered requests)"
)

# where the wrong loads land: a neighbour of the right skill, or somewhere unrelated?
misses = [
    (p["gold"], baseline[p["text"]]["loaded"][0])
    for p in POSITIVES
    if baseline[p["text"]]["loaded"] and baseline[p["text"]]["loaded"][0] != p["gold"]
]
same_category = sum(
    1
    for gold, got in misses
    if got in BY_NAME and BY_NAME[got]["category"] == BY_NAME[gold]["category"]
)
print(
    f"\nof {len(misses)} wrong first picks, {same_category} came from the right skill's own "
    f"category"
)
wrong loads    16.8%   (315 covered requests)
needless loads 9.8%   (173 uncovered requests)

of 36 wrong first picks, 10 came from the right skill's own category

Неверные загрузки попадают в собственную категорию правильного навыка гораздо чаще, чем это было бы случайно, поэтому трудность — различить несколько похожих друг на друга навыков. Агент уже ищет примерно в правильном месте.

Шаг 3: ранжируйте весь список

Один запрос несёт вопросы двух видов:

  • which — это вопрос Choice по всем 182 именам навыков, где описание из индекса служит критерием для каждого варианта (тот же текст, который получает и сам агент). Его вероятности — это ранжирование.
  • три вопроса Noul о запросе, напечатанные ниже; каждый по-своему спрашивает, требуется ли действие, а не объяснение. prose_suffices считает наоборот. Их среднее решает, подсказывать ли вообще, и ниже 0.30 не подсказывается ничего.

Оба уходят одним запросом, поэтому ранжирование и проверка стоят одного кругового рейса.

Напишите эти три вопроса так, чтобы они спрашивали, нужно ли действие. Вопрос о предмете не отделит объясни, что такое монада от запроса, которому нужен навык, — ведь и то и другое про софт.

Один вопрос Choice спокойно вмещает список такого размера. Если он в несколько раз больше, вы бы разбили его на куски и ранжировали каждый, а затем прогнали этот же шаг отбора по победителям.

CHOICE_INSTRUCTIONS = (
    "Which of these skills, if any, is the right one to load to help with the "
    "user's latest request?"
)
GATE_QUESTIONS = {
    "acts_on_user_system": (
        "Is the assistant being asked to act on the user's files, accounts, devices, "
        "or online services, rather than only to explain or advise?"
    ),
    "would_follow_documented_procedure": (
        "Would a careful expert answering this consult a specific documented procedure "
        "or set of commands, rather than answering from general understanding?"
    ),
    "prose_suffices": (
        "Could a knowledgeable generalist fully satisfy this request in prose, with "
        "no tools, no documentation, and no access to the user's files or accounts?"
    ),
}
INVERTED = {"prose_suffices"}  # a yes here points away from needing a skill

def build_state(request: str) -> dict:
    return {"request": request, "recent_context": ""}

@json_cache
def rank_wide(request: str) -> dict:
    """Request 1: rank all 182 skills, and score the request for whether a skill applies."""
    questions = {
        "which": Choice(
            instructions=CHOICE_INSTRUCTIONS,
            criteria={skill["name"]: skill["description"] for skill in ROSTER},
        )
    }
    for key, text in GATE_QUESTIONS.items():
        questions[f"gate::{key}"] = Noul(instructions=text)
    started = perf_counter()
    response = client.system_one(
        state=build_state(request), questions=questions, model=TYPESAFE_MODEL
    )
    ranked = sorted(
        response.answers["which"].probabilities.items(), key=lambda kv: -kv[1]
    )
    values = {
        key.removeprefix("gate::"): answer.noul
        for key, answer in response.answers.items()
        if key.startswith("gate::")
    }
    oriented = [(1.0 - v) if k in INVERTED else v for k, v in values.items()]
    return {
        "ranked": ranked[
            :12
        ],  # more than any shortlist needs, and keeps the cache small
        "gate": sum(oriented) / len(oriented),
        "values": values,
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

DEMO = [
    "Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so it syncs"
    " to my phone? Just write it up in whatever editor pops up.",
    "Can you put together a pitch deck skeleton (cover, situation overview, comps, precedent"
    " transactions, DCF, LBO) as a .pptx, using our firm-template.pptx for branding and"
    " footnoting each valuation number back to the cell it came from in the model?",
    "Post this announcement to my Mastodon account.",
]
for request in DEMO:
    wide = rank_wide(request)
    verdict = "suggest" if wide["gate"] >= GATE_THRESHOLD else "stay quiet"
    print(f'"{request[:78]}"')
    print(f"  needs a skill {wide['gate']:.2f} -> {verdict}   ({wide['seconds']}s)")
    for name, probability in wide["ranked"][:SHORTLIST]:
        print(f"    {probability:.3f}  {name:<38}{BY_NAME[name]['description']}")
    print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
  needs a skill 0.75 -> suggest   (0.31s)
    0.990  apple-notes                           Manage Apple Notes via memo CLI: create, search, edit.
    0.010  computer-use                          Drive the user's desktop in the background — clicking, ty...
    0.000  concept-diagrams                      Generate flat, minimal educational SVG visuals as HTML.

"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
  needs a skill 0.76 -> suggest   (0.16s)
    0.700  powerpoint                            Create, read, edit .pptx decks, slides, notes, templates.
    0.300  pptx-author                           Build PowerPoint decks headless with python-pptx.
    0.000  chroma                                Embedding database for RAG and semantic search.

"Post this announcement to my Mastodon account."
  needs a skill 0.78 -> suggest   (0.16s)
    0.550  xurl                                  X/Twitter via xurl CLI: raw post search, posting, DM, media.
    0.140  computer-use                          Drive the user's desktop in the background — clicking, ty...
    0.080  openhands                             Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).

Запрос про Notes.app однозначен, и его верхний вариант — правильный. Ничто из того, что может сделать ранжирование, не спасёт запрос про Mastodon: три вопроса говорят, что навык нужен, потому что публикация в аккаунт — это действие, а при наличии навыка для публикации в X и отсутствии навыка для Mastodon ближайший навык всё равно побеждает.

Остаётся дека. Оба лидера — навыки для .pptx, и на 60 символах широкий вопрос Choice ставит навык редактирования впереди навыка создания — для запроса о создании деки.

Шаг 4: переранжируйте тройку лучших

Три варианта оставляют место для полного описания плюс начала собственного SKILL.md каждого навыка, поэтому второй запрос ставит тот же вопрос перед лучшими данными:

  • which — вопрос Choice по шортлисту, где критерий каждого варианта — этот более длинный текст.
  • fits::{name} — один вопрос Noul на кандидата: делает ли этот навык конкретно то, о чём просит запрос? На каждый отвечают отдельно, поэтому все они могут вернуться низкими, а шортлист, у которого наивысший опустился ниже 0.30, отбрасывается целиком.
RERANK_INSTRUCTIONS = (
    "Exactly one of these skills is the right one to load for the user's latest "
    "request. Which one? Read what each actually does, not just its name."
)

def rerank_criteria(names: tuple[str, ...], excerpt: int) -> dict[str, str]:
    return {
        name: f"{BY_NAME[name]['description_full']} — {BY_NAME[name]['body'][:excerpt]}"
        for name in names
    }

def rerank_questions(names: tuple[str, ...], excerpt: int) -> dict:
    questions = {
        "which": Choice(
            instructions=RERANK_INSTRUCTIONS, criteria=rerank_criteria(names, excerpt)
        )
    }
    for name in names:
        questions[f"fits::{name}"] = Noul(
            instructions=(
                f"Does the skill '{name}' do the specific thing the user's request asks "
                f"for? It is described as: {BY_NAME[name]['description_full']}"
            )
        )
    return questions

@json_cache
def rerank(request: str, names: tuple[str, ...], excerpt: int) -> dict:
    """Request 2: the same Choice over a shortlist, plus one absolute noul per candidate."""
    started = perf_counter()
    response = client.system_one(
        state=build_state(request),
        questions=rerank_questions(names, excerpt),
        model=TYPESAFE_MODEL,
    )
    return {
        "winner": response.answers["which"].choice,
        "fits": {
            key.removeprefix("fits::"): answer.noul
            for key, answer in response.answers.items()
            if key.startswith("fits::")
        },
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

for request in DEMO:
    wide = rank_wide(request)
    if wide["gate"] < GATE_THRESHOLD:
        print(f'"{request[:78]}"\n  scored too low, nothing suggested\n')
        continue
    shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
    result = rerank(request, shortlist, EXCERPT_CHARS)
    best = max(result["fits"].values())
    verdict = result["winner"] if best >= FITS_THRESHOLD else "nothing fits"
    print(f'"{request[:78]}"')
    print(f"  was {shortlist[0]} -> {verdict}   ({result['seconds']}s)")
    for name in shortlist:
        print(f"    fits {result['fits'][name]:.2f}  {name}")
    print()
"Can you save this recipe as a new note in my 'Recipes' folder in Notes.app so "
  was apple-notes -> apple-notes   (0.12s)
    fits 0.60  apple-notes
    fits 0.54  computer-use
    fits 0.01  concept-diagrams

"Can you put together a pitch deck skeleton (cover, situation overview, comps, "
  was powerpoint -> pptx-author   (0.09s)
    fits 0.73  powerpoint
    fits 0.38  pptx-author
    fits 0.02  chroma

"Post this announcement to my Mastodon account."
  was xurl -> xurl   (0.09s)
    fits 0.56  xurl
    fits 0.38  computer-use
    fits 0.05  openhands

Два навыка для .pptx расходятся, как только каждый приносит свой собственный текст: запрос про деку переключается на навык создания.

Здесь fits-noul и Choice расходятся: noul оценивают навык редактирования выше, тогда как Choice выбирает навык создания. Они решают разные вещи. Choice решает, какой навык, а noul — говорить ли вообще хоть что-нибудь.

Запрос про Mastodon переживает обе проверки: его лучший fits-noul оказывается выше 0.30, поэтому рецепт предлагает навык для X для запроса про Mastodon. Большинство похожих запросов ловятся. Второй проход может отклонить только то, что ему передаёт широкое ранжирование, а здесь это были три близких промаха.

Функция ниже — весь рецепт: два запроса и два порога, и на выходе не более одного имени навыка.

Чтобы направить его на ваш собственный список, замените hermes_roster.json. Каждый вопрос выше читает из этого файла только name, description, description_full и body, и больше ничто не знает о Hermes.

def suggest(request: str) -> tuple[str, ...]:
    """At most one skill name for a request, or () for "nothing here applies"."""
    wide = rank_wide(request)
    if wide["gate"] < GATE_THRESHOLD:
        return ()
    shortlist = tuple(name for name, _ in wide["ranked"][:SHORTLIST])
    result = rerank(request, shortlist, EXCERPT_CHARS)
    if max(result["fits"].values()) < FITS_THRESHOLD:
        return ()
    return (result["winner"],)

def suggestion_block(names: tuple[str, ...]) -> str:
    """What gets appended after the roster, in the suggestion.

    This string is a measured input rather than prose: it goes to the agent, so it is part
    of every graded turn's cache key. Editing a word here silently invalidates the shipped
    results and costs a live re-run to restore them.
    """
    body = (
        f"Relevant to the current request: {', '.join(names)}. Ignore this if it does not "
        "fit what the user actually asked for."
        if names
        else "No skill in the roster appears relevant to this request."
    )
    return f"\n\n<skill_relevance>\n{body}\n</skill_relevance>"

print(suggestion_block(suggest(DEMO[1])))
print(suggestion_block(suggest(DEMO[2])))

<skill_relevance>
Relevant to the current request: pptx-author. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>

<skill_relevance>
Relevant to the current request: xurl. Ignore this if it does not fit what the user actually asked for.
</skill_relevance>

Шаг 5: измерьте подсказку

Каждый из 488 запросов идёт к агенту трижды, по одному измеряемому ходу. Прогоны различаются только тем, что сообщают агенту:

что попадает в системный промпт
агент один ничего
агент с подсказкой то, что вернула suggest()
агент, которому дали ответ имя покрывающего навыка или «ничего не подходит», если такого нет

Третье недостижимо; это потолок, с которым измеряются первые два.

Формулировка этой подсказки делает две вещи. Она говорит, что подсказку можно игнорировать, потому что более жёсткое давление повышает и согласие с неверными подсказками, а неверная хуже, чем никакая. А ход, которому нечего подсказать, всё равно шлёт предложение об этом; если не посылать ничего вовсе, инструкция списка «в сомнительных случаях загружай» осталась бы без противовеса.

texts = [request["text"] for request in REQUESTS]
with ThreadPoolExecutor(max_workers=WORKERS) as pool:  # up to 488 x 2 TypeSafe requests
    suggested = dict(zip(texts, pool.map(suggest, texts)))
WIDE = {text: rank_wide(text) for text in texts}  # all cache hits now; reused below

arms = {
    "baseline": {},
    "TypeSafe": {
        request["text"]: suggestion_block(suggested[request["text"]])
        for request in REQUESTS
    },
    "oracle": {
        request["text"]: suggestion_block((request["gold"],) if request["gold"] else ())
        for request in REQUESTS
    },
}
scores = {
    arm: summarise(run_arm(arm, suggestions)) for arm, suggestions in arms.items()
}

print(f"{'run':<10}{'wrong loads':>13}{'needless loads':>16}")
for arm, row in scores.items():
    print(f"{arm:<10}{row['wrong_load']:>13.1%}{row['needless_load']:>16.1%}")

def fewer(metric: str) -> str:
    """The plain ratio between the two arms' error rates."""
    return f"{scores['baseline'][metric] / scores['TypeSafe'][metric]:.1f}x fewer"

print(
    f"\nbaseline -> TypeSafe:  {fewer('wrong_load')} wrong loads, "
    f"{fewer('needless_load')} needless ones"
)
run         wrong loads  needless loads
baseline          16.8%            9.8%
TypeSafe           7.3%            4.0%
oracle             2.5%            1.2%

baseline -> TypeSafe:  2.3x fewer wrong loads, 2.4x fewer needless ones
moved = [
    (
        baseline[p["text"]]["loaded"][:1] == [p["gold"]],
        run_turn(AGENT_MODEL, "TypeSafe", p["text"], arms["TypeSafe"][p["text"]])[
            "loaded"
        ][:1]
        == [p["gold"]],
    )
    for p in POSITIVES
]
print(
    f"of {len(POSITIVES)} covered requests: {sum(not b and a for b, a in moved)} the suggestion "
    f"fixed, {sum(b and not a for b, a in moved)} it broke"
)
of 315 covered requests: 37 the suggestion fixed, 7 it broke

Подсказка исправляет гораздо больше запросов, чем ломает, но всё же ломает некоторые, которые агент делал верно сам. Уверенная неверная подсказка убедительнее, чем никакой подсказки, — такова цена того, чтобы поставить её перед ходом.

SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"

ARM_COLOR = {"baseline": BLUE, "TypeSafe": ORANGE, "oracle": MUTED}

def style(ax):
    ax.set_facecolor(SURFACE)
    for side in ("top", "right"):
        ax.spines[side].set_visible(False)
    for side in ("left", "bottom"):
        ax.spines[side].set_color(AXIS)
    ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
    ax.set_axisbelow(True)

panels = [
    ("wrong_load", f"wrong loads\n{len(POSITIVES)} covered requests"),
    ("needless_load", f"needless loads\n{len(NEGATIVES)} uncovered requests"),
]
names = list(scores)
fig, axes = plt.subplots(1, 2, figsize=(8.4, 3.6), facecolor=SURFACE)
for ax, (metric, title) in zip(axes, panels):
    style(ax)
    ax.grid(axis="y", color=GRID, linewidth=0.8)
    values = [scores[arm][metric] for arm in names]
    bars = ax.bar(
        names,
        values,
        0.58,
        color=[ARM_COLOR[arm] for arm in names],
        # the oracle is a ceiling, not a competitor: gray, and hatched so it never depends
        # on colour alone
        hatch=["", "", "///"],
        edgecolor=SURFACE,
        linewidth=1.2,
    )
    ax.bar_label(
        bars,
        labels=[f"{v:.1%}" for v in values],
        padding=3,
        color=INK2,
        fontsize=9,
    )
    ax.set_title(title, loc="left", color=INK2, fontsize=9.5)
    ax.set_ylim(0, max(values) * 1.28)
    ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))
    ax.set_ylabel("% of those requests - lower is better", color=INK2, fontsize=9)
fig.suptitle(
    f"Hermes' {len(ROSTER)}-skill roster, {len(REQUESTS)} requests, {AGENT_MODEL}",
    x=0.02,
    ha="left",
    color=INK,
    fontsize=11,
)
fig.tight_layout()
display(fig)
plt.close(fig)
вывод

Что показывают результаты

  • Неверные загрузки упали с 16.8% до 7.3%, а лишние — с 9.8% до 4.0%, а это большая часть разрыва между угадыванием по обрезанному индексу и получением готового ответа.
  • Некоторые запросы, которые агент делал верно сам, возвращаются неверными, как только подсказка приложена. Счётчики выше.

Копируйте эту схему, когда у вашего агента большой список: дешёвое ранжирование по всему, затем близкий взгляд на два-три. Любой из шагов может вернуться с пустыми руками.

Откройте в playground

Постройте ссылку на playground для запроса про деку из шага 4, используя полное описание и фрагмент тела каждого кандидата как его критерии.

demo_shortlist = tuple(name for name, _ in rank_wide(DEMO[1])["ranked"][:SHORTLIST])
playground_link = make_playground_link(
    build_state(DEMO[1]),
    rerank_questions(demo_shortlist, EXCERPT_CHARS),
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open the shortlist + questions in the TypeSafe playground]({playground_link})"
    )
)
Открыть шортлист и вопросы в TypeSafe Playground →

Что дальше

Та же схема встречается и в других местах: маршрутизация намерений — для маршрутизации к обработчику, а не к навыку, уверенность — для выбора двух порогов, и спекулятивный fan-out — для помещения всех вопросов в один запрос.