Open evaluation set / 56 rows

Can your system
stay inside the date?

A compact public benchmark for time-bounded historical answers: 26 honest unknowns and 30 grounded cases, derived from the rights-clean public pack.

Open github.com/datelock/benchmark
1

Take the rows unchanged

Use all 56 rows from the repository. Preserve each query, language, and as_of value. Do not tune on hidden knowledge or replace the historical cutoff.

2

Run your system

Produce one answer per row and retain the evidence your system used. UNKNOWN is a valid outcome; unsupported confidence is not.

3

Score reproducibly

Follow the benchmark repository's scoring instructions and report the exact revision, configuration, per-row outputs, and aggregate counts. Compare grounded and UNKNOWN rows separately; never convert a failed verification into credit.

Reference release

Rows
56
Grounded
30
Unknown
26
Pack prefix
5dd4c251