Evaluation API Reference

The evaluate() function allows you to benchmark Text-to-SQL model outputs against the LLMSQL gold queries and SQLite database. It prints metrics, logs mismatches, and saves detailed reports automatically.

Features

  • Evaluate model predictions from JSONL files or Python dicts.

  • Automatically download benchmark questions and SQLite DB if missing.

  • Prints mismatch summaries and supports configurable reporting.

  • Saves detailed JSON report with metrics, mismatches, timestamp, and input mode.

  • Optionally saves the results in the leaderboard run.yaml format.

Usage Examples

Evaluate from a JSONL file:

from llmsql.evaluation.evaluate import evaluate

report = evaluate("path_to_outputs.jsonl")
print(report)

Evaluate from a list of Python dicts:

predictions = [
    {"question_id": "1", "predicted_sql": "SELECT name FROM Table WHERE age > 30"},
    {"question_id": "2", "predicted_sql": "SELECT COUNT(*) FROM Table"},
]

report = evaluate(predictions)
print(report)

Using a persistent cache directory for benchmark downloads:

report = evaluate(
    "path_to_outputs.jsonl",
    workdir_path="./benchmark-cache",
)

Function Arguments

Argument

Description

outputs

Path to JSONL file or a list of prediction dicts (required).

workdir_path

Directory used to cache downloaded benchmark files. If omitted, a temporary directory is created automatically.

save_report

Path to save detailed JSON report. Defaults to “evaluation_results_{uuid}.json”.

show_mismatches

Print mismatches while evaluating. Default True.

max_mismatches

Maximum number of mismatches to display. Default 5.

model_name

Name of the evaluated model (e.g. Qwen/Qwen3-0.6B). Stored in the JSON report and the leaderboard YAML. Default None.

save_leaderboard_yaml

Optional path to also save the results in the leaderboard run.yaml format. Default None (not saved).

run_metadata

Optional dict deep-merged into the leaderboard YAML (model details, type, inference backend/arguments, device, …).

Input Format

The predictions should be in JSONL format:

{"question_id": "1", "predicted_sql": "SELECT name FROM Table WHERE age > 30"}
{"question_id": "2", "predicted_sql": "SELECT COUNT(*) FROM Table"}
{"question_id": "3", "predicted_sql": "SELECT * FROM Table WHERE active=1"}

Output Metrics

The function returns a dictionary with the following keys:

  • total – Total queries evaluated

  • matches – Queries where predicted SQL results match gold results

  • pred_none – Queries where the model returned NULL or no result

  • gold_none – Queries where the reference result was NULL or no result

  • sql_errors – Invalid SQL or execution errors

  • accuracy – Overall exact match accuracy

  • model_name – Name of the evaluated model (if provided)

  • version – LLMSQL benchmark version used for evaluation

  • mismatches – List of mismatched queries with details

  • timestamp – Evaluation timestamp

  • input_mode – How results were provided (“jsonl_path” or “dict_list”)

Report Saving

By default, a report is saved automatically as evaluation_results_{uuid}.json in the current directory. It contains metrics, mismatches, timestamp, and input mode. You can override this path using save_report.

Leaderboard Format

Pass save_leaderboard_yaml to additionally save the results in the same format as the run.yaml files in the leaderboard/ folder of the repository. The evaluation date, llmsql package version, benchmark version, OS name, Python version, execution accuracy, number of samples and outputs path are filled automatically; all other fields are null unless provided via run_metadata:

report = evaluate(
    "outputs.jsonl",
    model_name="Qwen/Qwen3-0.6B",
    save_leaderboard_yaml="run.yaml",
    run_metadata={
        "type": "open-source",
        "model": {"dtype": "bfloat16", "parameter_count": "0.6B"},
        "inference": {
            "backend": "vllm",
            "arguments": {"num_fewshots": 5, "temperature": 0.0},
        },
    },
)

From the CLI (--run-metadata accepts a YAML/JSON file with the same structure):

llmsql evaluate --outputs outputs.jsonl \
    --model-name Qwen/Qwen3-0.6B \
    --save-leaderboard-yaml run.yaml \
    --run-metadata metadata.yaml

—

LLMSQL Evaluation Module

Provides the evaluate() function to benchmark Text-to-SQL model outputs on the LLMSQL benchmark.

See the documentation for full usage details.

llmsql.evaluation.evaluate.evaluate(outputs: str | list[dict[int, str | int]], *, version: str = '2.0', workdir_path: str | None = None, save_report: str | None = None, show_mismatches: bool = True, max_mismatches: int = 5, model_name: str | None = None, save_leaderboard_yaml: str | None = None, run_metadata: dict[str, Any] | None = None) → dict[source]

Evaluate predicted SQL queries against the LLMSQL benchmark.

Parameters:
  • version – LLMSQL version

  • outputs – Either a JSONL file path or a list of dicts.

  • workdir_path – Directory to store downloaded benchmark files. If omitted, a temporary directory is created automatically.

  • save_report – Optional manual save path. If None → auto-generated.

  • show_mismatches – Print mismatches while evaluating.

  • max_mismatches – Max mismatches to print.

  • model_name – Name of the evaluated model (e.g. Qwen/Qwen3-0.6B). Stored in the JSON report and in the leaderboard YAML.

  • save_leaderboard_yaml – Optional path to additionally save the results in the leaderboard run.yaml format (see the leaderboard/ folder). If None, no YAML is written.

  • run_metadata – Optional dict deep-merged into the leaderboard YAML to fill fields that cannot be detected automatically, e.g. {"type": "open-source", "inference": {"backend": "vllm", "arguments": {"num_fewshots": 5}}}.

Returns:

Metrics and mismatches, plus coverage counters expected,

answered, missing and duplicates. accuracy is computed over the predictions, accuracy_over_benchmark counts unanswered questions as wrong.

Return type:

dict

Raises:

ValueError – if a prediction references a question_id that is not part of the benchmark.

—

💬 Made with ❤️ by the LLMSQL Team