트레이싱
한 호출의 모든 훅은 같은 PredictContext.run_id를 받으므로, 트레이서가 자체 장부를 유지하지
않고도 시작, 종료, 오류 이벤트(그리고 자신이 여는 모든 스팬)를 상관시킬 수 있습니다.
run_id
uuid4().hex문자열로, 공개 호출(predict_batch,system_one,Router.predict,ONNXAgent.system_one)마다 한 번 생성됩니다.Router.predict_batch는 대신 요청마다 하나를 생성하며, 그 요청에 대한Router.predict호출이 가졌을run_id입니다.- 그 호출의 모든 훅이 공유하며,
on_error와on_predict_end도 포함합니다. - 전역도 아니고 영속되지도 않습니다. 프로세스 내에서 호출 하나를 식별합니다. 시스템 간 상관을 위해 로그와 아웃바운드 페이로드에 넣으십시오.
- 호출마다 구별되므로 두 호출이 결코 충돌하지 않습니다.
call A: run_id=1a2b... start ──┐
├─ end
error ─┘
call B: run_id=9f8e... start ─── end
최소한의 트레이서
run_id에서 시작 시 연 스팬으로 가는 맵을 유지하고, 종료나 오류 시 닫으십시오.
import time
class Tracer:
def __init__(self):
self.spans = {}
def on_predict_start(self, ctx):
self.spans[ctx.run_id] = {
"model": ctx.model,
"started_at": ctx.started_at,
}
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
emit_span(
name="laya.predict",
run_id=ctx.run_id,
model=ctx.model,
duration_ms=ctx.elapsed_ms,
usage=ctx.usage,
ok=ctx.error is None,
)
def on_error(self, ctx):
# on_error runs before on_predict_end; leaving the span for on_predict_end is fine,
# or close it here if you prefer.
pass
agent = laya.load("convaiinnovations/laya", hooks=[Tracer()])
on_predict_end가 항상 실행되므로 스팬을 닫기에 자연스러운 자리이고, 실패 경로에서
ctx.error를 볼 수 있습니다.
스팬 수명
on_predict_start ──► open span (run_id, model, started_at)
│
├─ inference
│
on_error ──► record ctx.error on the span
│
on_predict_end ──► close span (elapsed_ms, usage, ok)
오류 스팬
on_predict_end는 ctx.error가 설정된 채 실패 경로에서 실행되므로, 닫는 지점 하나가 둘 다
처리합니다.
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
if ctx.error is not None:
span["status"] = "error"
span["error_type"] = type(ctx.error).__name__
span["error_message"] = str(ctx.error)
span["duration_ms"] = ctx.elapsed_ms
emit(span)
on_error만 구현한다면, 그것이 on_predict_end보다 먼저 실행된다는 것을 기억하십시오.
둘 다에서 스팬을 닫지 마십시오. 이중으로 집계됩니다.
OpenTelemetry
예제 examples/hooks/otel.py는 카운터와
히스토그램을 기록합니다. 실제 스팬에는 트레이서에서 OTel API를 구동하십시오. 훅은
동기이므로 동기 익스포터를 사용하십시오(또는 워커에서 큐에 넣고 내보내십시오).
from opentelemetry import trace
tracer = trace.get_tracer("laya")
class OTelHooks:
def __init__(self):
self.spans = {}
def on_predict_start(self, ctx):
span = tracer.start_span("laya.predict", attributes={"laya.run_id": ctx.run_id, "laya.model": ctx.model})
self.spans[ctx.run_id] = span
def on_predict_end(self, ctx):
span = self.spans.pop(ctx.run_id, None)
if span is None:
return
if ctx.usage:
span.set_attribute("laya.input_tokens", ctx.usage["input_tokens"])
if ctx.error is not None:
span.record_exception(ctx.error)
span.set_status(trace.Status(trace.StatusCode.ERROR))
span.end()
laya.load("convaiinnovations/laya", hooks=[OTelHooks()], hooks_raise=False)
트레이서 장애가 요청을 실패시키지 않도록 hooks_raise=False를 설정하십시오.
분산 전파
run_id는 일반 문자열이므로, 프로세스를 떠나는 모든 것에 넣으십시오. 로그 줄, 감사
서비스로 보내는 페이로드, 또는 결정이 다운스트림 호출을 유발할 때의 HTTP 헤더입니다.
def audit(ctx):
requests.post(
"https://audit.example/decisions",
json=record(ctx),
headers={"X-Laya-Run-Id": ctx.run_id},
timeout=2,
)
이런 블로킹 호출은 호출 스레드를 지연시킨다는 것을 기억하십시오. 동시에 서빙할 때는 대신 큐에 넣으십시오. 패턴을 참고하십시오.
중첩 호출
predict를 다시 호출하는 훅은 새 run_id로 새 호출을 시작합니다. 직접 연결하지 않으면
부모와 자식은 독립적입니다. 부모 id를 캡처해 함께 전달하십시오.
def enrich(ctx):
for state in ctx.states: # a hook sees every state of the call
child = enricher.predict(state, EXTRA_QUESTIONS)
record_child_span(parent_run_id=ctx.run_id, child_run_id=child.get("run_id"))
부모 run_id 하나, 상태당 자식 호출 하나입니다. ctx.states[0]을 읽는 본문은 배치의 첫
결정을 연결하고 나머지를 조용히 버립니다.
재귀를 방지하십시오(안티 패턴 참고). 가장 쉬운 가드는 훅이
없는 별도의 enricher 에이전트입니다.