문서

수명 주기

이 페이지는 각 진입점에 대한 정확한 이벤트 순서입니다. 다이어그램을 하나만 읽는다면 Router 것을 읽으십시오. 그것이 상위 집합입니다.

Agent.predict_batch

predict_batch가 유일한 구현이며, system_one과 predict는 상태 하나로 이를 호출합니다.

predict_batch(states, questions, batch_size=..., hooks=..., ...)
  │
  ├─ active  = installed hooks + per-call hooks         (installed first)
  ├─ ctx     = PredictContext(states, questions, model=self.model_id, agent=self)
  │
  ├─ try:
  │    │
  │    ├─ on_predict_start ─────────────────────────────┐
  │    │      a hook may:                                │
  │    │        • rewrite ctx.states / ctx.questions     │
  │    │        • set ctx.max_len / ctx.head_max_len     │
  │    │        • ctx.skip(results) ─────────────┐       │
  │    │        • raise (aborts; see errors)     │       │
  │    │                                         │      │
  │    ├─ if ctx.results is not None:  ◄─────────┘      │  cache hit
  │    │      skip tokenization and forward             │
  │    ├─ else:                                         │
  │    │      validate states is a list                 │
  │    │      for each batch chunk:                     │
  │    │        _encode_state ─► collate ─► _forward    │
  │    │      _decode_answers                           │
  │    │      ctx.results = [...]                       │
  │    │                                                │
  │    └─ (any failure here) ──► except BaseException:  │
  │              ctx.error = exc                        │
  │              on_error                               │
  │              re-raise                               │
  │                                                     │
  │   finally:                                          │
  │     ctx.elapsed_ms = now - ctx.started_at           │
  │     if ctx.results: ctx.usage = aggregate_usage(...)│
  │     on_predict_end ─────────────────────────────────┘
  │
  └─ return ctx.results

같은 ctx 객체가 시작, 오류, 종료를 통과하므로 run_id가 이들을 상관시키고, end 훅이 ctx.error를 읽을 수 있습니다.

Agent.system_one / predict

system_one(state, questions, hooks=..., ...)
  └─ predict_batch([state], questions, hooks=..., ...)[0]

따라서 system_one은 모든 훅과 같은 수명 주기를 상속하며, ctx.states == [state]입니다.

Router.predict

Router.predict(state, questions, model=..., hooks=..., on_predict_start=..., on_predict_end=...)
  │
  ├─ active = installed hooks + per-call hooks
  │
  ├─ route(state, questions, ..., hooks=per-call, hooks_raise=...)
  │    │
  │    ├─ _route(...)                      detect script / language / workflow
  │    ├─ on_route  ──► ctx.decision       a hook may replace the decision
  │    └─ return ctx.decision
  │
  ├─ load(decision["model"])
  │    │
  │    ├─ already resident? ──► return it
  │    ├─ else build Agent(...) ──► on_load   (after the Router lock is released)
  │    └─ evict LRU checkpoints ──► on_evict  (after the Router lock is released)
  │
  ├─ ctx = PredictContext(states=[state], questions, decision, model=decision.model,
  │                       agent=agent, router=self)
  ├─ try:
  │    ├─ on_predict_start
  │    ├─ if ctx.results is None:
  │    │      result = agent.system_one(ctx.states[0], ctx.questions)
  │    │        └─ the Agent's own hooks run here (start / forward / end)
  │    │      result["routing"] = decision
  │    │      ctx.results = [result]
  │    └─ else:
  │           for each cached result: result.setdefault("routing", decision)
  │    └─ (any failure) ──► except: on_error, re-raise
  │    └─ finally: elapsed_ms, usage, on_predict_end
  │
  └─ return ctx.results[0]

핵심 사항:

  • on_route는 모델이 로드되기 전에 실행되므로, 훅이 체크포인트를 고정해 다른 체크포인트 로드를 피할 수 있습니다.
  • Router 수준 예측 훅은 호출 전체를 감쌉니다. 이들은 Agent로 전달되지 않습니다. 자체 훅을 가진 연결된 Agent는 그 훅도 실행하며, 이는 예상된 동작입니다.
  • Router 수준의 ctx.skip()도 routing을 추가하므로 반환 형태가 안정적입니다.

Router.predict_batch

각 결과는 그 요청에 대해 predict가 반환하는 것이므로, Router 수준 예측 훅은 여기서도 요청마다 실행됩니다. 모든 요청이 자체 PredictContext, run_id, elapsed_ms를 가집니다.

Router.predict_batch(requests, batch_size=..., hooks=...)
  │
  ├─ active = installed hooks + per-call hooks          (installed first; None and [] add nothing)
  ├─ route_batch(requests, hooks=per-call) ──► on_route, once per request   (no checkpoint loaded yet)
  │
  └─ for each checkpoint, in order of first appearance:
       │
       ├─ load(checkpoint) ──► on_load / on_evict
       ├─ for each request of this checkpoint, in input order:
       │      ctx = PredictContext(states=[state], questions, decision, model, agent, router,
       │                           max_len=request.get("max_len"),
       │                           head_max_len=request.get("head_max_len"))
       │      on_predict_start       a hook may redact, rewrite, set a token budget or skip
       ├─ group the requests left to infer by (questions, ctx.max_len, ctx.head_max_len)
       │      agent.predict_batch(states, questions, ...)  ──► one shared forward pass per group
       │      result["routing"] = decision;  ctx.results = [result]
       ├─ ctx.usage, once per request
       ├─ (any failure) ──► for every started request, in reverse input order:
       │                    ctx.error = exc, on_error, on_predict_end;  then re-raise
       └─ on_predict_end, once per request of this checkpoint, in reverse input order

핵심 사항:

  • 한 체크포인트의 요청들에 대한 모든 시작 훅은 그들의 어떤 end 훅보다도 먼저 실행됩니다. 순전파를 공유하기 때문입니다. 따라서 on_predict_end에서 채우는 캐시는 같은 체크포인트 그룹 안의 중복 상태를 제공할 수 없고, 호출을 넘어서는 제공할 수 있습니다.
  • 같은 이유로 요청들은 시작한 순서의 역순으로 끝나므로, 시작에서 무언가를 설정하고 end에서 되돌리는 훅(contextvars 값, OpenTelemetry context.attach / detach)이 처음 찾은 값으로 풀립니다.
  • ctx.states, ctx.questions 또는 토큰 예산을 교체하는 시작 훅은 자기 요청만 바꿉니다. 요청은 시작 훅이 실행된 뒤에 순전파를 위해 그룹화되기 때문입니다. 공유 질문 딕셔너리를 제자리에서 변경하는 것은 다르고, predict가 하는 일도 아닙니다. 그룹의 어떤 것도 모든 시작 훅이 실행될 때까지 추론되지 않으므로, 그 변경은 그 딕셔너리를 공유하는 모든 요청 (시작 훅이 더 일찍 실행된 요청 포함)과 호출자에게도 도달합니다. 대신 ctx.questions에 새 딕셔너리를 할당하십시오.
  • 체크포인트 이름은 그룹화 전에 해석되므로, 별칭("ml")으로 요청을 고정하는 on_route 훅은 그 체크포인트의 순전파를 공유하고, ctx.model은 해석된 이름입니다.
  • 요청은 자체 max_len / head_max_len을 지닐 수 있습니다. 이는 predict가 호출 인자로 받는 토큰 예산의 요청별 형태입니다. ctx.max_len을 설정하는 시작 훅이 이를 덮어씁니다. 훅이 컨텍스트가 만들어진 뒤에 실행되기 때문입니다. 다른 예산을 요청하는 요청들은 순전파를 공유할 수 없으므로, 예산이 섞인 배치는 예산마다 agent.predict_batch 호출을 하나씩 만듭니다.
  • 시작된 모든 요청은 정확히 하나의 on_predict_end를 받으며, 다른 요청의 end 훅이 예외를 던져도 마찬가지입니다. 그런 첫 오류는 그것들이 모두 실행된 뒤에 발생합니다.
  • 체크포인트 그룹이 실패하면, 시작된 모든 요청이 그 예외로 실패합니다. 각각 ctx.error가 그것으로 설정된 on_error를, 그다음 on_predict_end를 받습니다. 여기에는 캐시 히트와 질문 그룹이 이미 실행된 요청도 포함됩니다. 호출자가 그들 중 어느 것에 대해서도 결과 없이 예외를 받기 때문이며, 이는 ctx.error가 다른 요청의 실패일 수 있음을 뜻합니다(한 요청에 대해 예외를 던지는 시작 훅은 자기 그룹을 실패시킵니다). 이미 끝난 체크포인트 그룹의 요청은 [router.predict(...) for ...]의 이전 호출들이 그랬을 것처럼 결과와 함께 끝났습니다.

모델 수명 주기

on_load는 체크포인트가 만들어질 때, on_evict는 하나가 해제될 때 실행됩니다. 둘 다 Router 내부 잠금이 풀린 뒤에 실행되므로, 훅이 Router를 안전하게 다시 호출할 수 있습니다.

load("multilingual")
  │
  ├─ [lock]
  │    build Agent(...)          (seconds: download + weights)
  │    register in _agents / _order
  │    evict LRU if over max_loaded ──► evicted = ["english"]
  ├─ [unlock]
  ├─ on_evict("english")
  └─ on_load("multilingual")

unload("english")
  ├─ [lock] remove from _agents / _order
  ├─ [unlock]
  └─ on_evict("english")

attach(name, agent)는 기존 에이전트를 등록하며, 체크포인트가 만들어지지 않았으므로 on_load를 실행하지 않습니다.

skip을 사용한 캐싱

on_predict_start
  ├─ cache hit?  ctx.skip([cached_result])
  │     └─ forward pass skipped
  │     └─ on_predict_end still runs
  │     └─ Router adds `routing` if missing
  └─ cache miss? nothing
        └─ forward pass runs
        └─ on_predict_end can store the result

동작하는 캐시는 examples/hooks/cache.py를 참고하십시오.

빈 입력

감사가 모든 호출을 보도록 훅은 여전히 실행됩니다.

입력 종료 시 ctx.results
predict_batch([]) []
predict_batch(states, {})(질문 없음) 상태당 빈 답변 페이로드 하나
system_one(state, {}) 빈 답변 페이로드 하나

이 경우에는 토큰화나 순전파가 일어나지 않지만, on_predict_start와 on_predict_end는 실행됩니다.

순서 규칙

  1. 설치된 훅은 항상 호출별 훅보다 먼저 실행됩니다.
  2. 리스트 안에서는 리스트 순서대로 실행됩니다.
  3. 한 이벤트에 대해, 그것을 구현한 모든 훅이 그 순서대로 다음 이벤트 전에 실행됩니다.
  4. 실패 경로에서 on_error는 on_predict_end보다 먼저 실행됩니다.
  5. 하나의 load가 해제와 생성을 모두 할 때 on_evict는 on_load보다 먼저 실행됩니다.
installed: [A, B]   per-call: [C]
on_predict_start: A, B, C
on_predict_end:   A, B, C

동시성

Agent와 Router는 여러 스레드에서 호출해도 안전합니다. 각 호출이 자체 PredictContext를 만들므로 컨텍스트가 요청 간에 새지 않습니다. 유일한 공유 상태는 훅 객체 자체이므로, 스레드 안전하지 않은 훅은 자체 상태를 보호하거나 hooks_concurrent=False로 설치해야 합니다.

hooks_concurrent=True (default)      hooks_concurrent=False
  thread 1 ─┐                          thread 1 ─┐
  thread 2 ─┼─ hooks run in parallel   thread 2 ─┼─ one hook at a time
  thread 3 ─┘                          thread 3 ─┘   (RLock)

hooks_concurrent=False는 각 훅 호출을 직렬화하는 것이지, 호출 전체를 직렬화하는 것이 아닙니다. 두 호출은 여전히 이벤트 사이에서 교차할 수 있습니다. 재진입 잠금을 사용하므로, 훅이 교착 없이 같은 Agent/Router를 다시 호출할 수 있습니다.

함께 보기

  • 오류: 훅이 예외를 던지면 이벤트별로 무슨 일이 일어나는지.
  • 패턴과 안티 패턴: 수명 주기를 잘 활용하는 방법.