Evaluation API Reference¶
The evaluate() function allows you to benchmark Text-to-SQL model outputs against the LLMSQL gold queries and SQLite database. It prints metrics, logs mismatches, and saves detailed reports automatically.
Features¶
Evaluate model predictions from JSONL files or Python dicts.
Automatically download benchmark questions and SQLite DB if missing.
Prints mismatch summaries and supports configurable reporting.
Saves detailed JSON report with metrics, mismatches, timestamp, and input mode.
Optionally saves the results in the leaderboard
run.yamlformat.
Usage Examples¶
Evaluate from a JSONL file:
from llmsql.evaluation.evaluate import evaluate
report = evaluate("path_to_outputs.jsonl")
print(report)
Evaluate from a list of Python dicts:
predictions = [
{"question_id": "1", "predicted_sql": "SELECT name FROM Table WHERE age > 30"},
{"question_id": "2", "predicted_sql": "SELECT COUNT(*) FROM Table"},
]
report = evaluate(predictions)
print(report)
Using a persistent cache directory for benchmark downloads:
report = evaluate(
"path_to_outputs.jsonl",
workdir_path="./benchmark-cache",
)
Function Arguments¶
Argument |
Description |
|---|---|
outputs |
Path to JSONL file or a list of prediction dicts (required). |
workdir_path |
Directory used to cache downloaded benchmark files. If omitted, a temporary directory is created automatically. |
save_report |
Path to save detailed JSON report. Defaults to “evaluation_results_{uuid}.json”. |
show_mismatches |
Print mismatches while evaluating. Default True. |
max_mismatches |
Maximum number of mismatches to display. Default 5. |
model_name |
Name of the evaluated model (e.g. |
save_leaderboard_yaml |
Optional path to also save the results in the leaderboard |
run_metadata |
Optional dict deep-merged into the leaderboard YAML (model details, |
Input Format¶
The predictions should be in JSONL format:
{"question_id": "1", "predicted_sql": "SELECT name FROM Table WHERE age > 30"}
{"question_id": "2", "predicted_sql": "SELECT COUNT(*) FROM Table"}
{"question_id": "3", "predicted_sql": "SELECT * FROM Table WHERE active=1"}
Output Metrics¶
The function returns a dictionary with the following keys:
total – Total queries evaluated
matches – Queries where predicted SQL results match gold results
pred_none – Queries where the model returned NULL or no result
gold_none – Queries where the reference result was NULL or no result
sql_errors – Invalid SQL or execution errors
accuracy – Overall exact match accuracy
model_name – Name of the evaluated model (if provided)
version – LLMSQL benchmark version used for evaluation
mismatches – List of mismatched queries with details
timestamp – Evaluation timestamp
input_mode – How results were provided (“jsonl_path” or “dict_list”)
Report Saving¶
By default, a report is saved automatically as evaluation_results_{uuid}.json in the current directory. It contains metrics, mismatches, timestamp, and input mode. You can override this path using save_report.
Leaderboard Format¶
Pass save_leaderboard_yaml to additionally save the results in the same format as the
run.yaml files in the leaderboard/ folder of the repository. The evaluation date,
llmsql package version, benchmark version, OS name, Python version, execution accuracy,
number of samples and outputs path are filled automatically; all other fields are null
unless provided via run_metadata:
report = evaluate(
"outputs.jsonl",
model_name="Qwen/Qwen3-0.6B",
save_leaderboard_yaml="run.yaml",
run_metadata={
"type": "open-source",
"model": {"dtype": "bfloat16", "parameter_count": "0.6B"},
"inference": {
"backend": "vllm",
"arguments": {"num_fewshots": 5, "temperature": 0.0},
},
},
)
From the CLI (--run-metadata accepts a YAML/JSON file with the same structure):
llmsql evaluate --outputs outputs.jsonl \
--model-name Qwen/Qwen3-0.6B \
--save-leaderboard-yaml run.yaml \
--run-metadata metadata.yaml
—
LLMSQL Evaluation Module¶
Provides the evaluate() function to benchmark Text-to-SQL model outputs on the LLMSQL benchmark.
See the documentation for full usage details.
- llmsql.evaluation.evaluate.evaluate(outputs: str | list[dict[int, str | int]], *, version: str = '2.0', workdir_path: str | None = None, save_report: str | None = None, show_mismatches: bool = True, max_mismatches: int = 5, model_name: str | None = None, save_leaderboard_yaml: str | None = None, run_metadata: dict[str, Any] | None = None) dict[source]¶
Evaluate predicted SQL queries against the LLMSQL benchmark.
- Parameters:
version – LLMSQL version
outputs – Either a JSONL file path or a list of dicts.
workdir_path – Directory to store downloaded benchmark files. If omitted, a temporary directory is created automatically.
save_report – Optional manual save path. If None → auto-generated.
show_mismatches – Print mismatches while evaluating.
max_mismatches – Max mismatches to print.
model_name – Name of the evaluated model (e.g.
Qwen/Qwen3-0.6B). Stored in the JSON report and in the leaderboard YAML.save_leaderboard_yaml – Optional path to additionally save the results in the leaderboard
run.yamlformat (see theleaderboard/folder). If None, no YAML is written.run_metadata – Optional dict deep-merged into the leaderboard YAML to fill fields that cannot be detected automatically, e.g.
{"type": "open-source", "inference": {"backend": "vllm", "arguments": {"num_fewshots": 5}}}.
- Returns:
- Metrics and mismatches, plus coverage counters
expected, answered,missingandduplicates.accuracyis computed over the predictions,accuracy_over_benchmarkcounts unanswered questions as wrong.
- Metrics and mismatches, plus coverage counters
- Return type:
dict
- Raises:
ValueError – if a prediction references a
question_idthat is not part of the benchmark.
—