トレーシング
トレーシング
1 回の呼び出しのすべてのフックは同じ PredictContext.run_id を受け取るので、tracer は自前の帳簿付けをせずに開始・終了・エラーの各イベント(および開いた span)を相関できます。
run_id
uuid4().hexの文字列で、公開呼び出しごとに 1 つ作られます(predict_batch、system_one、Router.predict、ONNXAgent.system_one)。Router.predict_batchは代わりにリクエストごとに 1 つ作ります。そのリクエストに対するRouter.predict呼び出しが持ったであろうrun_idです。- その呼び出しのすべてのフックが共有します。
on_errorとon_predict_endも含みます。 - グローバルではなく、永続化もされません。プロセス内で呼び出しを識別します。システムをまたいで相関させるには、ログと外向きのペイロードに入れてください。
- 呼び出しごとに異なるので、2 つの呼び出しが衝突することはありません。
call A: run_id=1a2b... start ──┐
├─ end
error ─┘
call B: run_id=9f8e... start ─── end
最小限の tracer
run_id から、start で開いた span へのマップを保ち、end か error でそれを閉じます。
import time
class Tracer:
def __init__(self):
self.spans = {}
def on_predict_start(self, ctx):
self.spans[ctx.run_id] = {
"model": ctx.model,
"started_at": ctx.started_at,
}
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
emit_span(
name="laya.predict",
run_id=ctx.run_id,
model=ctx.model,
duration_ms=ctx.elapsed_ms,
usage=ctx.usage,
ok=ctx.error is None,
)
def on_error(self, ctx):
# on_error runs before on_predict_end; leaving the span for on_predict_end is fine,
# or close it here if you prefer.
pass
agent = laya.load("convaiinnovations/laya", hooks=[Tracer()])
on_predict_end は必ず走るので、span を閉じる自然な場所であり、失敗経路では ctx.error も見られます。
span の寿命
on_predict_start ──► open span (run_id, model, started_at)
│
├─ inference
│
on_error ──► record ctx.error on the span
│
on_predict_end ──► close span (elapsed_ms, usage, ok)
エラー span
on_predict_end は失敗経路でも ctx.error が設定された状態で走るので、閉じる場所は 1 つで両方を扱えます。
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
if ctx.error is not None:
span["status"] = "error"
span["error_type"] = type(ctx.error).__name__
span["error_message"] = str(ctx.error)
span["duration_ms"] = ctx.elapsed_ms
emit(span)
on_error だけを実装する場合は、それが on_predict_end より先に発火することを忘れないでください。両方で span を閉じると二重計上になります。
OpenTelemetry
例 examples/hooks/otel.py はカウンタとヒストグラムを記録します。本物の span が必要なら、tracer から OTel API を駆動します。フックは同期なので、同期の exporter を使ってください(または enqueue し、worker から export します)。
from opentelemetry import trace
tracer = trace.get_tracer("laya")
class OTelHooks:
def __init__(self):
self.spans = {}
def on_predict_start(self, ctx):
span = tracer.start_span("laya.predict", attributes={"laya.run_id": ctx.run_id, "laya.model": ctx.model})
self.spans[ctx.run_id] = span
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
if ctx.usage:
span.set_attribute("laya.input_tokens", ctx.usage["input_tokens"])
if ctx.error is not None:
span.record_exception(ctx.error)
span.set_status(trace.Status(trace.StatusCode.ERROR))
span.end()
laya.load("convaiinnovations/laya", hooks=[OTelHooks()], hooks_raise=False)
hooks_raise=False にすると、tracer の障害がリクエストを失敗させることが決してありません。
分散伝播
run_id は普通の文字列なので、プロセスを出るものに入れてください。ログ行、監査サービスへ送るペイロード、あるいは意思決定が下流呼び出しを引き起こすときの HTTP ヘッダです。
def audit(ctx):
requests.post(
"https://audit.example/decisions",
json=record(ctx),
headers={"X-Laya-Run-Id": ctx.run_id},
timeout=2,
)
このようなブロッキング呼び出しが呼び出しスレッドを止めることを忘れないでください。並行して提供するときは、代わりに enqueue してください。パターンを参照。
ネストした呼び出し
再び predict を呼ぶフックは、新しい run_id を持つ新しい呼び出しを始めます。自分で結びつけない限り、親と子は独立です。親の id を取り出して渡してください。
def enrich(ctx):
for state in ctx.states: # a hook sees every state of the call
child = enricher.predict(state, EXTRA_QUESTIONS)
record_child_span(parent_run_id=ctx.run_id, child_run_id=child.get("run_id"))
親の run_id が 1 つ、state ごとに子呼び出しが 1 つです。ctx.states[0] を読むだけの実装はバッチの最初の意思決定だけを結びつけ、残りを黙って捨てます。
再帰を防いでください(アンチパターンを参照)。最も簡単な防壁は、フックを持たない別の enricher エージェントです。
関連項目
- API リファレンス:完全な
PredictContext。 - パターンとアンチパターン:ノンブロッキングなトレーシング。