빠른 시작
Ollaya는 오픈 의사결정 모델을 자신의 머신에서 실행합니다. 모델에 상태(메시지, 이메일, 티켓, JSON 객체)와 몇 개의 타입이 지정된 질문을 주면, 한 번의 순전파에서 캘리브레이션된 확률이 담긴 타입 지정 답을 돌려줍니다.
1. 설치
Linux와 macOS에서는:
curl -fsSL https://ollaya.dev/install.sh | sh
Windows에서는 PowerShell로:
irm https://ollaya.dev/install.ps1 | iex
또는 데스크톱 앱을 받으십시오. 같은 명령줄과 서버가 함께 들어 있습니다.
스크립트는 GitHub에서 최신 릴리스를 내려받아 sha256을 검사합니다. Linux와 Windows에서는 NVIDIA GPU를 찾으면 CUDA 라이브러리도 가져옵니다. Linux에서 systemd가 돌고 root 권한이 있으면 127.0.0.1:11435에서 API를 제공하는 서비스를 설정합니다. 요구 사항과 Docker 이미지는 다운로드를 참고하십시오.
2. 모델 실행하기
ollaya run winnow:e4b "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent refund ███████████████░ 0.91
is_urgent yes ███████████████░ 0.92
frustration 2.89 / 3 very angry or using stron… ██████████████░░ 0.86
refund_requested yes ████████████████ 0.99
churn_risk yes ████████████████ 0.99
ollaya run은 서버가 실행 중이 아니면 시작하고, 처음 사용할 때 모델을 내려받아 로드합니다. winnow:e4b는 권장 모델입니다. 타입 지정 의사결정에서 정확도 0.722로 TypeSafe의 Jev(0.738)에 가깝고, RTX 4090에서 이 다섯 질문에 89 ms가 걸립니다. 다운로드는 8 GB입니다.
NVIDIA GPU가 없습니까? CPU에서 1초의 몇 분의 1 만에 답하는 laya로 시작하십시오. 이것은 router입니다. 영어 텍스트를 laya:en으로, 다른 언어(예: 터키어)를 laya:multilingual로 보내며, laya를 받으면 둘 다 받아집니다. Models가 모든 모델의 정확도와 속도를 비교합니다.
--preset NAME은 내장 질문 세트를 사용합니다:triage,email,guard,moderation,router,agent. 지정하지 않으면 자체 질문이 없는 모델은triage에 답합니다.--verbose는 각 선택지의 확률, 라우팅 결정, 소요 시간을 추가로 출력합니다.--format json은 전체 API 응답을 출력합니다.- 상태가 없으면
ollaya run은 파이프된 stdin을 읽거나 터미널에서 프롬프트를 엽니다.
3. 직접 질문 만들기
질문을 파일에 씁니다:
{
"topic": {
"type": "choice",
"instructions": "What is this message about?",
"criteria": {
"billing": "Payments, invoices and refunds",
"access": "Login, passwords and permissions",
"other": "Anything else"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Needs attention this week", "Needs attention today"]
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
}
}
ollaya run winnow:e4b --questions questions.json "Hi, I cannot log in since this morning and I have a demo at 3pm."
또는 파일을 건너뛰고 curl -d처럼 JSON을 인라인으로 전달합니다({로 시작하는 값은 파일 경로가 아니라 JSON입니다):
ollaya run winnow:e4b --questions '{"angry":{"type":"noul","instructions":"Is the customer angry?"}}' "Hi, I cannot log in since this morning and I have a demo at 3pm."
또는 같은 질문을 API로 보냅니다:
curl http://localhost:11435/api/decide -d '{
"model": "winnow:e4b",
"state": "Hi, I cannot log in since this morning and I have a demo at 3pm.",
"questions": {
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
모든 답은 타입이 지정되어 돌아옵니다. choice는 선택지마다 확률을, score는 기대 단계를, noul은 그 명제가 성립할 확률을 줍니다. API 레퍼런스를 참고하십시오.
4. 기존 TypeSafe 클라이언트 사용하기
Ollaya는 TypeSafe의 API도 제공합니다. 공식 TypeSafe Python SDK는 변경 없이 동작합니다:
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value; the SDK needs one
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
TypeSafe 호환성을 참고하십시오.
5. 질문을 모델에 굽기
Modelfile은 질문 세트를 이름으로 실행할 수 있는 모델로 바꿉니다:
FROM laya
QUESTIONS ./questions.json
DESCRIPTION Support inbox triage
ollaya create inbox -f Modelfile
ollaya run inbox "Hi, I cannot log in since this morning and I have a demo at 3pm."