LlamaIndex 集成
LlamaIndex 集成
Laya 为 LlamaIndex 的 RAG 流水线、RouterQueryEngine 和工具选择提供低于 35ms 的非自回归
决策组件(单问题延迟:Tesla T4 GPU 上用 laya-multilingual 测得 32.8 ms,用 laya 测得
39.5 ms;CPU 上为 193–464 ms):
LayaSingleSelector:低于 35ms 的单选选择器,替代RouterQueryEngine的LLMSingleSelector。LayaMultiSelector:多选选择器,替代LLMMultiSelector,用于横跨多个数据源的复合查询。LayaQueryRouter:独立的查询分发器,把进来的请求直接路由到目标查询引擎或可调用对象。
同时支持本地进程内推理(Agent 或 Router)和远程 HTTP 推理(对接你自己的
laya-serve 实例),边缘客户端不需要装 PyTorch。
安装
pip install "laya[llamaindex]"
1. 用 RouterQueryEngine 做单选路由
在 LlamaIndex 里,RouterQueryEngine 用一个选择器决定该由哪个底层查询引擎或工具来回答
问题。自回归的 LLM 选择器(LLMSingleSelector)要花 1,000–2,000 ms 生成文本。LayaSingleSelector
用 ~33 ms 评估候选工具,不做 token 生成:
from llama_index.core.query_engine import RouterQueryEngine
from llama_index.core.tools import QueryEngineTool, ToolMetadata
from laya.integrations.llamaindex import LayaSingleSelector
# Define query engine tools
docs_tool = QueryEngineTool(
query_engine=vector_index.as_query_engine(),
metadata=ToolMetadata(
name="vector_documentation",
description="Semantic search over technical user documentation and API guides.",
),
)
sql_tool = QueryEngineTool(
query_engine=sql_index.as_query_engine(),
metadata=ToolMetadata(
name="sql_database",
description="Structured SQL database containing customer accounts, billing, and orders.",
),
)
# Initialize Laya sub-35ms selector with confidence fallback
selector = LayaSingleSelector(
confidence_threshold=0.80, # If confidence < 0.80, fall back to index 0
fallback_index=0,
)
router_engine = RouterQueryEngine(
selector=selector,
query_engine_tools=[docs_tool, sql_tool],
)
response = router_engine.query("What is the shipping address for order #4912?")
print(response)
2. 复合查询的多选
对于需要在多个索引之间综合的查询(比如把文档规格与事务数据库记录做对比),LayaMultiSelector
评估候选项的相关性,返回多个被选中的工具:
from laya.integrations.llamaindex import LayaMultiSelector
multi_selector = LayaMultiSelector(
probability_threshold=0.25, # Select all tools with probability >= 0.25
max_outputs=2,
)
tools = [docs_tool.metadata, sql_tool.metadata, summary_tool.metadata]
result = multi_selector.select(
tools,
"How does the database security policy compare with our published compliance guide?"
)
for sel in result.selections:
print(f"Tool: {tools[sel.index].name} | {sel.reason}")
3. 用 LayaQueryRouter 直接分发查询
要做直接路由、又不想背上 RouterQueryEngine 的开销,LayaQueryRouter 可以把查询直接路由到
用字典注册的引擎:
from laya.integrations.llamaindex import LayaQueryRouter
router = LayaQueryRouter(
query_engines={
"vector": vector_query_engine,
"sql": sql_query_engine,
"summary": summary_query_engine,
},
descriptions={
"vector": "Semantic search over product documentation and guides",
"sql": "Structured SQL queries for user accounts and transactions",
"summary": "Quarterly reports and high-level business summaries",
},
confidence_threshold=0.75,
fallback_key="vector",
)
# Route and execute in one call:
response = router.query("How many active subscriptions were renewed in Q3?")
print(response)
同步的 query() 和异步的 aquery() 都支持。
4. 置信度阈值门控
和 Laya 的 LangChain 集成一样,LayaSingleSelector 和 LayaQueryRouter 读取校准过的
answer_confidence(max(p)):
- 自动回退: 指定
fallback_index(或fallback_key),把不确定的查询无缝转到一个安全的 默认引擎。 - 严格防护: 在
LayaSingleSelector上设置raise_on_low_confidence=True,输入含义不清时 抛出LayaLowConfidenceError,让调用方升级处理。
5. 远程 HTTP 部署
适用于 serverless RAG、边缘环境或没有本地 GPU 的环境:
from laya.integrations.llamaindex import LayaSingleSelector
selector = LayaSingleSelector(
base_url="http://laya-serve.internal:8080",
confidence_threshold=0.85,
fallback_index=0,
)
远程客户端用 Python 标准库的 urllib,不引入任何重依赖,避免跨源转发凭据,并符合
/v1/systemone 规范。
6. 逐调用的决策控制
LayaSingleSelector、LayaMultiSelector 和 LayaQueryRouter 接受和核心 API 一样的逐调用参数:
两个 token 预算(max_len、head_max_len)和五个预测钩子参数(hooks、on_predict_start、
on_predict_end、hooks_raise、hooks_timeout)。它们是按选择器设置的,所以可以让一个宽度大
的路由步骤多留些空间,而流水线的其余部分仍保持 checkpoint 的默认值。
一个 choice 问题的各个选项共享 checkpoint 的选项预算 —— head_max_len,在 laya 上是 192 个
token —— 而每个候选都会把自己的名字和描述贡献进去,所以超过大约 20 个工具之后,这些描述开始以
同样的文本到达模型。
selector = LayaSingleSelector(
instructions="Which tool or query engine is best suited to answer this query?",
max_len=1024, # total window
head_max_len=512, # tokens shared by the option prompt
)
result = selector.select(tools, query) # tools: 59 descriptions
在 laya 上测量(Apple 芯片,每个查询一次前向传播,按选中的工具计分),用一份 59 个工具的名单,
它是从 MASSIVE 英文意图标签构建的、每个标签一句话语,所以真值是精确的。每个单元格是 59 个查询里
有多少个命中了它自己的工具;两次重复得到了相同的计数。
| 59 工具名单 | 默认预算 | max_len=1024, head_max_len=384 |
…, head_max_len=512 |
|---|---|---|---|
| 命中自己工具的查询 | 2/59 | 8/59 | 15/59 |
| 每个查询的中位 ms | 160 | 172 | 184 |
这里主张的不是绝对准确率:这个 checkpoint 不是 MASSIVE 分类器,而 59 个相似的标签是一种压力形状。 主张的是方向和代价 —— 一份会被默认预算压到近乎为零的名单变得可读了,而在这种规模下,加宽窗口 几乎不花时间。选项少于约 20 个时,标签本来就已经装得下,加宽反而可能把答案带偏,这就是两个参数 都按选择器选择加入的原因。那条实测的悬崖见 LangChain 集成。
钩子只在本地路径上运行。 一个带 base_url 和 hooks=[...] 的选择器会抛出 ValueError,
而不是报告一个其缓存从未运行过的成功 —— 钩子是一个在 predict 内部运行的 Python 可调用对象,
没有任何线上格式能携带它。请把钩子装在运行推理的那个进程里。两个预算确实会随请求体传到远程节点,
上限是该节点的 LAYA_MAX_TOKEN_BUDGET;更大的值会以 422 返回。