Score
Score는 서술적인 순서형 레벨에 따라 콘텐츠를 평가하기 위한 System One 질문 유형입니다. 답변에는 score, 레벨별 확률, 그리고 신뢰도가 포함됩니다.
답변이 단계로 설명할 수 있는 스펙트럼 위의 위치일 때 Score를 사용하십시오. 예를 들어 버그가 얼마나 심각한지, 고객이 얼마나 만족하는지, 지원자가 Python 경험을 얼마나 가졌는지 등입니다. 답변이 서로 순서가 없는 고정된 선택지 집합 중 하나라면 Choice를 사용하십시오. yes 또는 no라면 Noul을 사용하십시오. 질문 유형 선택하기에서 세 가지를 모두 비교합니다.
Score 답변은 score에 담긴, 여러분의 레벨을 따라가는 위치이며 두 레벨 사이에 떨어질 수 있습니다. 모델은 또한 probabilities에 모든 레벨의 확률을, 답변에 대한 confidence 값을 반환합니다.
예시 점수 질문
How severe is the reported issue?
상태 (평가할 내용)
The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.
답
각 단계의 확률
신뢰도
점수: 1.43
점수와 신뢰도는 어떻게 계산되나
점수:
각 단계의 번호에 그 확률을 곱한 뒤 모두 더합니다:
0 × 0 + 1 × 0.57 + 2 × 0.43 ≈ 1.43
신뢰도
TypeSafe는 확률이 각 단계에 흩어진 정도로 신뢰도를 계산합니다. 한 단계에 모두 몰려 있으면 1.0이고, 분포가 고를수록 낮아집니다.
How formal is this outfit based on the description?
상태 (평가할 내용)
A navy blazer over a plain white T-shirt, dark jeans, and clean leather loafers. No tie.
답
각 단계의 확률
신뢰도
점수: 1.86
점수와 신뢰도는 어떻게 계산되나
점수:
각 단계의 번호에 그 확률을 곱한 뒤 모두 더합니다:
0 × 0 + 1 × 0.14 + 2 × 0.86 + 3 × 0 + 4 × 0 ≈ 1.86
신뢰도
TypeSafe는 확률이 각 단계에 흩어진 정도로 신뢰도를 계산합니다. 한 단계에 모두 몰려 있으면 1.0이고, 분포가 고를수록 낮아집니다.
How relevant is this candidate's experience to the job posting?
상태 (평가할 내용)
Job posting: Senior backend engineer building Python APIs and PostgreSQL services. Candidate: Three years building Django REST APIs with PostgreSQL, preceded by two years in frontend JavaScript. Has owned small services but has not led a backend team.
답
각 단계의 확률
신뢰도
점수: 2.52
점수와 신뢰도는 어떻게 계산되나
점수:
각 단계의 번호에 그 확률을 곱한 뒤 모두 더합니다:
0 × 0 + 1 × 0 + 2 × 0.48 + 3 × 0.52 ≈ 2.52
신뢰도
TypeSafe는 확률이 각 단계에 흩어진 정도로 신뢰도를 계산합니다. 한 단계에 모두 몰려 있으면 1.0이고, 분포가 고를수록 낮아집니다.
How frustrated is the customer?
상태 (평가할 내용)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
답
각 단계의 확률
신뢰도
점수: 1.26
점수와 신뢰도는 어떻게 계산되나
점수:
각 단계의 번호에 그 확률을 곱한 뒤 모두 더합니다:
0 × 0 + 1 × 0.74 + 2 × 0.26 ≈ 1.26
신뢰도
TypeSafe는 확률이 각 단계에 흩어진 정도로 신뢰도를 계산합니다. 한 단계에 모두 몰려 있으면 1.0이고, 분포가 고를수록 낮아집니다.
How much does the report give an engineer to work with?
상태 (평가할 내용)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
답
각 단계의 확률
신뢰도
점수: 3.00
점수와 신뢰도는 어떻게 계산되나
점수:
각 단계의 번호에 그 확률을 곱한 뒤 모두 더합니다:
0 × 0 + 1 × 0 + 2 × 0 + 3 × 1 ≈ 3.00
신뢰도
TypeSafe는 확률이 각 단계에 흩어진 정도로 신뢰도를 계산합니다. 한 단계에 모두 몰려 있으면 1.0이고, 분포가 고를수록 낮아집니다.
각 단계 앞의 숫자는 위치이며, 레벨에서 설명합니다.
요청 구조
TypeSafe API로 보내는 POST 요청 본문은 다른 질문 유형과 마찬가지로 세 개의 최상위 필드를 가집니다. 평가할 콘텐츠인 state, model, 그리고 questions입니다. 각 Score 질문에는 다음 필드가 있습니다.
type: 항상"score"입니다.instructions: 모델이 답하는 질문입니다. 무엇을 평가하는지입니다.criteria: 척도의 낮은 쪽에서 높은 쪽으로 정렬된 레벨 설명의 배열입니다. 최소 두 개의 레벨이 있어야 하며, API는 최대 10개까지 받습니다.
아래는 상태가 버그 보고이고 질문은 버그가 얼마나 심각한지인 요청입니다.
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}질문 id는 여러분이 고르며, 여기서는 bug_severity입니다. 이 id는 모델에 전송되지 않습니다. 답변은 같은 id 아래에 반환됩니다.
레벨
criteria의 각 항목은 하나의 레벨입니다. 가능한 답변의 스펙트럼 위의 한 지점을 말로 설명한 것입니다. 레벨의 번호는 criteria 배열에서의 위치이며 0부터 시작하므로, 위의 세 항목은 레벨 0, 1, 2입니다. 배열의 순서가 곧 번호입니다.
모델은 설명만 받고 그 밖에는 아무것도 받지 않으며, 각 레벨은 상태에 대해 개별적으로 판단됩니다.
응답의 score는 레벨 스펙트럼 위의 위치입니다. 세 레벨 척도의 경우 0에서 2까지이며, 두 레벨 사이에 떨어질 수 있습니다.
저희 클라이언트 SDK는 타입이 지정된 질문을 제공합니다. Python에서 같은 질문은 Score입니다.
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score)
System One 모델을 호출하려면 system_one 메서드 또는 https://api.typesafe.ai/v1/systemone 엔드포인트를 사용하십시오. model 필드가 어느 모델이 요청을 처리할지 선택합니다. 코드의 어디에서 호출할지는 TypeSafe로 구축하는 방법에서 다룹니다.
저희 클라이언트 SDK 중 하나를 사용하거나 TypeSafe API를 직접 호출하십시오. 코딩 에이전트가 통합을 대신 작성한다면, 요청과 응답 형태를 알 수 있도록 먼저 TypeSafe 에이전트 스킬을 설치하십시오.
응답 구조
응답에는 요청의 id 아래에 질문당 answers 항목이 하나씩 있습니다. 다음은 위 예시 요청에 대한 응답입니다.
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 332,
"output_tokens": 18
}
}
각 Score 답변에는 다섯 개의 값이 있습니다.
type: TypeSafe 질문의 유형입니다.probabilities: 레벨 번호를 문자열로 키로 한 각 레벨의 확률입니다. 모든 값의 합은 1입니다.score: 레벨 번호 수직선 위의 위치로, 0에서 최상위 레벨 번호(여기서는 2)까지입니다. 각 레벨 번호에 그 확률을 곱해 더한 값입니다. 0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43.legend: 각 레벨 번호를 그 설명에 다시 매핑한 것입니다.confidence:probabilities가 얼마나 퍼져 있는지로 계산된 0에서 1 사이의 숫자입니다. 한 레벨에 단일 봉우리가 있으면 높은 신뢰도입니다. 확률이 여러 레벨에 퍼져 있으면 낮은 신뢰도입니다.
1.43이라는 score는 모델이 레벨 1과 2 사이에서 갈라져 레벨 1로 기운다는 뜻입니다. 이는 보고와 일치합니다. export가 고장 났고, Chrome으로 바꾸는 것이 대부분 고객에게는 우회책이지만 Safari만 쓰는 고객에게는 그렇지 않습니다. 모델은 “workaround exists”에 0.57을, “no workaround”에 0.43을 부여하고, 갈라져 있으므로 신뢰도는 0.35입니다.
Python SDK에서 ScoreAnswer는 score, confidence, probabilities, legend를 타입이 지정된 필드로 가집니다. SDK는 probabilities와 legend를 문자열이 아니라 정수 레벨로 키를 정합니다.
Score 읽기
입력이 다를 때 score가 어떻게 변하는지 살펴보겠습니다. 예를 들어 위 요청의 질문과 그 레벨을 사용하면 다음과 같습니다.
"How severe is the reported issue?"
→ 0: Cosmetic; no impact to functionality
→ 1: Broken or degraded feature, but workaround exists
→ 2: Blocking issue; no workaround exists
서로 다른 버그 보고가 score를 어떻게 바꾸는지 볼 수 있습니다.
probabilities | |||||
|---|---|---|---|---|---|
| 상태 | score | confidence | 레벨 0 | 레벨 1 | 레벨 2 |
| The export button is misaligned by a few pixels on the settings page. | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
| The PDF export button does nothing when clicked. I can still export to CSV and convert it myself, but that takes ages. | 1.0 | 1.0 | 0.0 | 1.0 | 0.0 |
| Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. | 1.11 | 0.84 | 0.0 | 0.89 | 0.11 |
| The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari. | 1.43 | 0.35 | 0.0 | 0.57 | 0.43 |
| Nobody on our team can log in since this morning. We get a 500 error on every attempt. | 2.0 | 1.0 | 0.0 | 0.0 | 1.0 |
이 예시들에서 신뢰도 1.0은 반환된 분포가 모든 확률을 한 레벨에 둔다는 뜻입니다. 이는 모델의 답변을 설명하는 것이지, 그 답변이 정확하다는 보장이 아닙니다.
score는 레벨 번호의 확률 가중 평균입니다. 세 번째와 네 번째 예시에서는 확률이 레벨 1과 2 사이로 갈라집니다. 레벨 2에 더 많은 가중치가 실리면 score가 올라갑니다. 우회책이 없는 고객의 비율을 측정하는 것이 아닙니다.
서로 다른 분포가 같은 score를 낼 수 있습니다. score 1.0은 모든 확률이 레벨 1에 있다는 뜻일 수도 있고, 레벨 0과 2에 절반씩 있다는 뜻일 수도 있습니다. 이를 구분하려면 score와 함께 probabilities와 confidence를 읽으십시오.
소수 score는 위치입니다. 보고를 심각도로 순위 매기는 데 사용하거나, 코드에 하나의 결과가 필요할 때 가장 가까운 레벨로 반올림할 수 있습니다. 저희 엔터티 정렬 쿡북은 결정을 내리기 위해 가장 가까운 레벨로 반올림하는 예를 보여줍니다.
Score에서 낮은 신뢰도는 보통 세 가지 중 하나를 뜻합니다. 이 상태에 대해 레벨이 겹치거나, 질문이 둘 이상을 측정하거나, 상태가 자리매김할 만큼 충분히 말하지 않는 경우입니다. 저희 신뢰도 문서는 코드에서 이를 사용하는 방법을 다룹니다.
좋은 레벨 작성하기
정도가 아니라 상황을 설명하십시오. “Broken or degraded feature, but workaround exists”는 모델이 상태를 대조할 무언가를 줍니다. “Moderately severe”는 그렇지 않습니다. 구체적인 설명은 모델이 레벨을 구분하는 데 도움이 될 수 있습니다. 알려진 예시로 답변을 확인하십시오. 더 높은 신뢰도만으로는 어떤 설명이 더 낫다는 것이 드러나지 않습니다.
모든 레벨은 개별적으로 평가됩니다. 모델은 레벨의 번호나 이웃을 보지 못하므로 “이전 레벨보다 나쁨”은 아무 의미가 없고, 설명이나 instructions에 있는 숫자도 도움이 되지 않습니다. 다음은 위 표의 잘못 정렬된 버튼 보고에 대해 레벨이 숫자만 있을 때 벌어지는 일입니다.
instructions: "Rate severity from 0 to 2, where 2 is worst"
criteria: ["0", "1", "2"]
→ score 0.55, confidence 0.33, probabilities 0: 0.45, 1: 0.55, 2: 0.0
같은 보고를 세 개의 서술적 레벨로 하면 신뢰도 1.0으로 0.0을 받습니다. 숫자만 있으면 모델이 대조할 것이 없어 확률을 0과 1 사이로 나눕니다.
구분되게 설명할 수 있는 만큼 레벨을 사용하십시오. 최대 10개까지입니다. 세 개면 충분합니다. 구분되게 설명할 수 없는 레벨은 추가하지 마십시오.
각 Score 질문은 하나의 차원으로 유지하십시오. 설명이 “punctual and smart and experienced”라고 말한다면, 질문이 세 가지를 측정하는 것이고, 한쪽은 높고 다른 쪽은 낮은 입력은 자리매김할 수 없습니다. 신뢰도가 떨어지고 score의 의미가 줄어듭니다. 다음 절에서 보여주듯이 한 가지당 Score 질문 하나로 나누고 코드에서 결합하십시오.
척도의 최상단에 다르게 행동해야 하는 드문 극단 사례가 있다면 그것에 고유한 레벨을 주십시오. “very angry”에서 끝나는 감정 척도는 “abusive or threatening”을 추가할 수 있습니다. 그 레벨이 없으면 두 메시지가 모두 최상단 근처의 score를 받을 수 있습니다. score만으로는 둘을 구분하지 못할 수 있습니다.
중간이 전혀 없고 답변이 몇 개의 이산적 범주 중 하나라면, 대신 Choice를 사용하거나 질문을 여러 Noul 질문으로 나누십시오. 레벨을 자체 데이터로 테스트하는 것이 중요합니다. 같은 척도의 두 가지 표현이 여러분의 데이터에서 다르게 작동할 수 있습니다.
복잡한 판단을 여러 Score 질문으로 나누기
여러 가지에 달려 있는 복잡한 판단은 한 가지당 Score 질문 하나로 나누는 것이 가장 좋습니다. 그런 다음 TypeSafe에서 반환된 Score들을 코드에서 결합해 판단을 내릴 수 있습니다. 어떤 Score 질문은 다른 것보다 더 중요할 수 있으므로, 각 Score 질문에 상대적 중요도에 따른 가중치를 부여하십시오. 가중치는 여러분의 것입니다. 결합된 결과가 팀이 내릴 결정과 맞지 않으면 코드에서 바꾸고 다시 실행하십시오. Score 질문은 하나의 요청으로 보내십시오. 병렬로 평가됩니다. 질문을 추가해도 응답 시간은 거의 변하지 않고 약간의 추가 질문 토큰만 소모합니다. 여러 질문을 함께 하기를 참조하십시오.
아래 요청은 위 표의 스피너 티켓에 맥락을 좀 더한 것입니다. 세 개의 Score 질문을 합니다. 버그가 얼마나 심각한지, 고객이 얼마나 불만인지, 보고서가 엔지니어에게 얼마나 많은 정보를 주는지입니다.
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave"
]
},
"report_quality": {
"type": "score",
"instructions": "How much does the report give an engineer to work with?",
"criteria": [
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment"
]
}
}
}TypeSafe의 응답:
{
"model": "jev-1.13.0",
"answers": {
"severity": {
"type": "score",
"score": 1.24,
"confidence": 0.64,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.76,
"2": 0.24
}
},
"frustration": {
"type": "score",
"score": 1.28,
"confidence": 0.58,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language or threatening to leave"
},
"probabilities": {
"0": 0.0,
"1": 0.72,
"2": 0.28
}
},
"report_quality": {
"type": "score",
"score": 3.0,
"confidence": 1.0,
"legend": {
"0": "No detail; just says something is broken",
"1": "Names the feature but no steps or environment",
"2": "Steps to reproduce or environment, but not both",
"3": "Steps to reproduce and environment"
},
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
},
"usage": {
"input_tokens": 468,
"output_tokens": 43
}
}
각 질문은 티켓에 대해 개별적으로 답변되고 score를 받습니다.
severity는 신뢰도 0.64로 1.24입니다. 처음 예시와 같은 해석입니다. export가 고장 났고 일부는 우회책이 있습니다.frustration은 신뢰도 0.58로 1.28입니다. 표현은 정중하지만 “third time”과 “I’m done”이 score의 일부를 최상위 레벨로 옮기므로, 모델은 “frustrated but civil”과 “very angry” 사이에 0.72와 0.28을 나눕니다. 이 티켓에서는 두 레벨이 겹치므로 신뢰도가 중간입니다.report_quality는 신뢰도 1.0으로 3.0입니다. 단계와 브라우저 버전이 모두 명시되어 있습니다.
세 척도의 길이가 다르므로 결합하기 전에 각 score를 정규화하십시오. 네 레벨 척도는 0에서 3을, 세 레벨 척도는 0에서 2를 반환하므로, 한쪽의 최상위 score가 다른 쪽의 최상위 score보다 큽니다. 각 score를 최상위 레벨 번호인 len(criteria) - 1로 나누어 모든 score를 0에서 1로 맞추십시오. 그러면 가중치가 말하는 대로 의미하게 됩니다. severity에 0.6, frustration에 0.3을 주면 severity가 두 배로 셈해집니다.
아래 TypeSafe Python SDK 코드는 세 질문을 하고, 각 score를 정규화하고, 예시 우선순위 계산을 사용해 결합합니다.
from typesafe_sdk import Score, TypeSafeClient
TRIAGE_QUESTIONS = {
"severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave",
],
),
"report_quality": Score(
instructions="How much does the report give an engineer to work with?",
criteria=[
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment",
],
),
}
def normalized(answers, question_id: str) -> float:
"""Put a score on 0 to 1 by dividing by its top level number."""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(
state=ticket,
questions=TRIAGE_QUESTIONS,
)
answers = response.answers
severity = normalized(answers, "severity")
frustration = normalized(answers, "frustration")
report_quality = normalized(answers, "report_quality")
# A detailed report helps an engineer investigate, so it raises priority a little.
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
위 예시 응답의 경우 정규화된 score는 severity 0.62, frustration 0.64, report quality 1.0입니다. 우선순위는 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664이며, 반올림하면 0.66입니다.
가중치는 여러분의 코드에 있으므로 숫자가 어떻게 만들어지는지 정확히 볼 수 있고, 순위가 팀이 할 것과 맞지 않을 때 바꿀 수 있습니다. 나중에 더 많은 Score 질문이 필요하면 TRIAGE_QUESTIONS에 추가하십시오. 요청 횟수는 하나로 유지됩니다. 복잡한 판단을 개별 Score로 나눈 다음 코드에서 가중치로 결합하는 이 기법을 복합 스코어링 패턴이라고 합니다.
구조화된 레벨 설명
각 레벨에 기본 텍스트 설명으로 시작하십시오. 모델이 분명하다고 생각하는 입력에서 계속 두 이웃 레벨 사이에 점수를 매길 때, 각 레벨을 문자열 대신 객체로 주고, 그 레벨이 무엇을 다루는지에 대한 필드와 몇 가지 예시 상황이 담긴 필드를 두십시오. 같은 필드 이름을 모든 레벨에 사용해 모델이 같은 것끼리 비교할 수 있게 하십시오.
아래 요청은 앞서 사용한 스피너 티켓이지만 각 레벨에 예시가 붙어 있습니다.
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
{
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
{
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
{
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
]
}
}
}응답:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.09,
"confidence": 0.87,
"legend": {
"0": {
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
"1": {
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
"2": {
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
},
"probabilities": {
"0": 0.0,
"1": 0.91,
"2": 0.09
}
}
},
"usage": {
"input_tokens": 379,
"output_tokens": 18
}
}
순수 문자열로 이 티켓은 신뢰도 0.84로 1.11을 받았습니다. 예시를 붙이면 신뢰도 0.87로 1.09를 받는데, 순수 문자열이 이미 잘 자리매김했기 때문에 작은 변화입니다. 순수 문자열이 모델을 갈라지게 할 때 효과가 더 큽니다. 다음 표에서 보여줍니다.
예시는 모델을 이끌며, 여러분의 실제 입력처럼 보일 때만 도움이 됩니다. 아래 표는 처음의 Safari 보고에 세 가지 레벨 객체 집합을 적용한 것입니다.
| 레벨 설명 | score |
confidence |
|---|---|---|
| 순수 문자열: 예시가 있는 객체 없음 | 1.43 | 0.35 |
| 유용한 예시가 담긴 examples 배열 추가: “export fails in one browser but works in another” | 1.03 | 0.96 |
| 브라우저와 무관한 예시가 담긴 examples 배열 추가: “search fails, but browsing categories still works” | 1.43 | 0.35 |
이 비교에서 일치하는 예시는 거의 모든 확률을 한 레벨에 집중시킵니다. 무관한 예시는 순수 문자열과 같은 결과를 냅니다. 더 높은 신뢰도가 어느 답이 정확한지 확립하지는 않습니다. 예상 레벨을 알고 있는 예시를 고르고, 수정한 설명을 별도의 입력으로 테스트한 뒤 유지하십시오.