文件導航

鏈路追蹤

一次呼叫的每個鉤子都收到同一個 PredictContext.run_id,所以一個 tracer 可以關聯起始、結束和 錯誤事件(以及它開啟的任何 span),而不用自己記賬。

run_id

  • 一個 uuid4().hex 字串,每次公開呼叫建立一次(predict_batch、system_one、 Router.predict、ONNXAgent.system_one)。Router.predict_batch 改為每個請求建立一個,即 該請求的一次 Router.predict 呼叫本會有的那個 run_id。
  • 被那次呼叫的每個鉤子共享,包括 on_error 和 on_predict_end。
  • 不是全域性的,也不持久化:它在程序內標識一次呼叫。把它放進你的日誌和出站載荷裡,以便跨系統關聯。
  • 每次呼叫各不相同,所以兩次呼叫永不衝突。
call A:  run_id=1a2b...  start ──┐
                                 ├─ end
                          error ─┘
call B:  run_id=9f8e...  start ─── end

一個最小 tracer

維護一個從 run_id 到你 start 時開啟的 span 的對映,在 end 或 error 時關閉它。

import time

class Tracer:
    def __init__(self):
        self.spans = {}

    def on_predict_start(self, ctx):
        self.spans[ctx.run_id] = {
            "model": ctx.model,
            "started_at": ctx.started_at,
        }

    def on_predict_end(self, ctx):
        span = self.spans.pop(ctx.run_id, None)
        if span is None:
            return
        emit_span(
            name="laya.predict",
            run_id=ctx.run_id,
            model=ctx.model,
            duration_ms=ctx.elapsed_ms,
            usage=ctx.usage,
            ok=ctx.error is None,
        )

    def on_error(self, ctx):
        # on_error runs before on_predict_end; leaving the span for on_predict_end is fine,
        # or close it here if you prefer.
        pass

agent = laya.load("convaiinnovations/laya", hooks=[Tracer()])

因為 on_predict_end 總會執行,它是關閉 span 的自然位置,而且它在失敗路徑上能看到 ctx.error。

span 的生命週期

on_predict_start ──► open span (run_id, model, started_at)
      │
      ├─ inference
      │
on_error ──► record ctx.error on the span
      │
on_predict_end ──► close span (elapsed_ms, usage, ok)

錯誤 span

on_predict_end 在失敗路徑上執行,且 ctx.error 被設定,所以一個關閉位置就能處理兩種情況:

def on_predict_end(self, ctx):
    span = self.spans.pop(ctx.run_id, None)
    if span is None:
        return
    if ctx.error is not None:
        span["status"] = "error"
        span["error_type"] = type(ctx.error).__name__
        span["error_message"] = str(ctx.error)
    span["duration_ms"] = ctx.elapsed_ms
    emit(span)

如果你只實現 on_error,記住它在 on_predict_end 之前觸發;不要兩處都關閉 span,否則會重複計數。

OpenTelemetry

示例 examples/hooks/otel.py 記錄計數器和一張直方圖。要真正的 span,就從 tracer 裡驅動 OTel API。鉤子是同步的,所以用同步的 exporter(或者入隊,再從一個 worker 匯出)。

from opentelemetry import trace

tracer = trace.get_tracer("laya")

class OTelHooks:
    def __init__(self):
        self.spans = {}

    def on_predict_start(self, ctx):
        span = tracer.start_span("laya.predict", attributes={"laya.run_id": ctx.run_id, "laya.model": ctx.model})
        self.spans[ctx.run_id] = span

    def on_predict_end(self, ctx):
        span = self.spans.pop(ctx.run_id, None)
        if span is None:
            return
        if ctx.usage:
            span.set_attribute("laya.input_tokens", ctx.usage["input_tokens"])
        if ctx.error is not None:
            span.record_exception(ctx.error)
            span.set_status(trace.Status(trace.StatusCode.ERROR))
        span.end()

laya.load("convaiinnovations/laya", hooks=[OTelHooks()], hooks_raise=False)

設定 hooks_raise=False,這樣一個 tracer 出故障永遠不會讓請求失敗。

分散式傳播

run_id 是一個普通字串,所以把它放進任何離開程序的東西里:日誌行、發給審計服務的載荷,或者一次 決策觸發下游呼叫時的 HTTP 頭。

def audit(ctx):
    requests.post(
        "https://audit.example/decisions",
        json=record(ctx),
        headers={"X-Laya-Run-Id": ctx.run_id},
        timeout=2,
    )

記住像這樣的阻塞呼叫會拖住呼叫執行緒;在併發服務時改成入隊。見 模式。

巢狀呼叫

一個再次呼叫 predict 的鉤子會啟動一次新呼叫,帶一個新的 run_id。父與子是獨立的,除非你自己 把它們連起來。捕獲父 id 並傳下去:

def enrich(ctx):
    for state in ctx.states:                     # a hook sees every state of the call
        child = enricher.predict(state, EXTRA_QUESTIONS)
        record_child_span(parent_run_id=ctx.run_id, child_run_id=child.get("run_id"))

一個父 run_id,每個狀態一次子呼叫。一個讀 ctx.states[0] 的寫法只關聯一個批次裡的第一個決策, 並靜默丟掉其餘的。

防範遞迴(見反模式);最簡單的防線是一個單獨的、沒有鉤子的 enricher agent。

另見