Scorer
Scores text from 0 to 100 against a criteria, with a panel of domain experts.
NOTE
Outputs on this page are illustrative. LLM responses vary between runs and models. The shape of the result stays the same, but the exact numbers will differ from what you see here.
How it works
Scorer assembles a jury and averages its verdicts:
- If you don't supply juries,
ScorerasksLister.with_jurieswhich expert roles suit the text and criteria. - Each jury scores the text independently, 0-100, with its own reasoning.
Scorercombines those scores into afinal_scorewith summary reasoning.
Picking the juries from the content is what makes the score meaningful. Medical text gets scored by a cardiologist and a medical writer instead of by a generic reviewer.
Basic usage
content = "Added rate limiting with a sliding window algorithm, including unit tests and performance benchmarks"
criteria = "Evaluate technical quality, completeness, and engineering best practices"
result = ActiveGenie::Scorer.call(content, criteria)
result.data
# => 91
result.reasoning
# => "All three reviewers rate the implementation highly, citing the appropriate
# algorithm choice and the presence of both tests and benchmarks."Supply your own juries when you know which perspectives matter:
juries = ["Cardiologist", "Clinical Researcher", "Medical Writer"]
result = ActiveGenie::Scorer.call(
"Patient shows 17% improvement in cardiac ejection fraction following a 6-week therapy protocol",
"Evaluate clinical accuracy and reporting quality",
juries: juries
)Interface
.call(text, criteria, juries: [], config: {})
Scores the text. Alias for .by_jury_bench.
| Parameter | Type | Description |
|---|---|---|
text | String | The content to score. |
criteria | String | What the juries should evaluate against. |
juries | Array<String> | Optional. Expert roles to use. Accepts both keyword (juries: [...]) and positional argument. |
config | Hash | Per-call configuration overrides. See Configuration. |
juries can be passed as a keyword argument (juries: [...]) matching Ranker, or as a 3rd positional argument for backward compatibility.
.by_jury_bench(text, criteria, juries: [], config: {})
Identical to .call. Use it when you want the strategy named explicitly at the call site.
Return value
Returns an ActiveGenie::Result.
| Accessor | Type | Contents |
|---|---|---|
data | Numeric | The final score, 0-100. A single number, not a hash. |
reasoning | String | Summary reasoning behind the final score. |
metadata | Hash | Per-jury scores and reasoning, keyed by strings. |
IMPORTANT
result.data is the score itself. The per-jury breakdown lives in metadata, with string keys derived from each jury's role.
result = ActiveGenie::Scorer.call("The sky is blue.", "Factual accuracy")
result.data
# => 100
result.metadata
# => {
# "meteorologists_reasoning" => "Accurate description of Rayleigh scattering under clear conditions.",
# "meteorologists_score" => 100,
# "physicists_reasoning" => "Correct as a plain statement of observed colour.",
# "physicists_score" => 100,
# "general_public_reasoning" => "Universally understood and correct.",
# "general_public_score" => 100,
# "final_score" => 100,
# "final_reasoning" => "All reviewers agree the statement is factually correct."
# }Jury key names come from the roles chosen for that call. When you don't supply juries yourself, the roles vary between runs, and so do the metadata keys. Supply juries explicitly if you need stable keys.
Configuration
Scorer performs best with gemini-3.5-flash-lite and gpt-5.6-luna (achieving up to 100% pass rates in live benchmark testing).
ActiveGenie::Scorer.call(text, criteria, config: { llm: { model: 'gemini-3.5-flash-lite' } })See Configuration for the full set of options, and Observability & errors for failure handling.
Cost
- Supplying juries costs one LLM call.
- Leaving juries empty costs two calls: one to select the jury, one to score.
Passing juries explicitly halves the cost and makes the output keys deterministic. If you score many items against the same criteria, select the jury once and reuse it:
juries = ActiveGenie::Lister.with_juries(sample_text, criteria).data
documents.map { |doc| ActiveGenie::Scorer.call(doc, criteria, juries) }Tips
- Make the criteria specific. "Evaluate quality" produces arbitrary numbers. "Evaluate medical accuracy, clarity, and clinical relevance" produces defensible ones.
- Scores are comparative, not absolute. A 78 means little on its own. Use
Scorerto rank or threshold items evaluated under identical criteria rather than as a certificate against an absolute standard. - Supply juries for consistency. Generated juries differ between runs, which moves the score. Fixed juries make scores comparable across a batch.
- Don't over-trust small gaps. An 84 and an 87 are not reliably distinguishable. Use
Comparatorwhen you need a confident head-to-head decision. - Thresholding works better than exact values. A rule like "≥ 70 passes" holds up across runs, while a rule that expects exactly 85 will break.
Examples
Content quality
ActiveGenie::Scorer.call(
article_body,
"Evaluate clarity, factual accuracy, and usefulness to a beginner audience",
["Editor", "Subject Matter Expert", "Beginner Reader"]
)Code review
ActiveGenie::Scorer.call(
diff,
"Evaluate correctness, test coverage, and adherence to SOLID principles",
["Senior Software Engineer", "Security Engineer"]
)Compliance
result = ActiveGenie::Scorer.call(
marketing_copy,
"Evaluate compliance with advertising standards and absence of unsubstantiated claims",
["Compliance Officer", "Legal Counsel"]
)
flag_for_review! if result.data < 70