# Jordan Applied AI Benchmark: example pack v0.1

**Status: not run.** Six paired examples, 12 prompts total, in English and Arabic. This is a public synthetic protocol sample, not the proposed future 36-pair dataset. No model/API execution, performance result, independent bilingual review, national ranking or training credential is established.

Each organization, person implied by a handoff, product, order, source snippet and business record is fictional and original. Jordan context comes from JOD, Amman-local time and Arabic/English work. The examples do not represent Jordan’s population or industry. No healthcare, real-client or actual government claim is included.

| Family | Pair ID | Focus |
| --- | --- | --- |
| Evidence-based research | `jaai-ebr-001` | Separate source support, planned targets and missing outcomes |
| Business decision support | `jaai-dec-001` | Respect two budget caps and calculate hypothetical savings |
| Bilingual work communication | `jaai-com-001` | Preserve a Jordanian Arabic handoff in an English draft |
| Automation planning | `jaai-aut-001` | Order dependencies and retain missing-input/approval boundaries |
| AI product scoping | `jaai-scope-001` | Bound a staff-reviewed reply assistant and acceptance behavior |
| Commerce data quality | `jaai-cat-001` | Detect nine catalog issues without inventing repairs |

## Files and format

- `tasks.json`: stable metadata and 12 tasks. Task IDs end in `-en` or `-ar`; each pair has identical structured inputs, output schema, reference assertions, scoring rule and abstention conditions. Only the instruction language changes. The communication pair requests an English update in both languages.
- `task-schema.json`: JSON Schema Draft 2020-12 for the pack. It enforces 12 tasks, two per family and six per locale. A separate integrity check must enforce unique IDs, one EN/AR task per pair and pair parity.
- `results.json`: explicit `not_run` status with empty `runs` and `results`; zero counts mean no attempts, not zero performance.
- `scorer.py`: optional offline machine scorer for a supplied local JSON answer. It does not execute models, rate language-review assertions or record results.
- `provenance.txt`: original synthetic authorship and existing public terms.
- `manifest.json`: SHA-256 and byte length of every other file; no self-hash.

Every task supplies `task_id`, `pair_id`, `locale`, `family`, `prompt`, `inputs`, `allowed_tools`, `required_output`, `reference_assertions`, `scoring_rule` and `abstention_conditions`. All allowed-tool lists are empty. Source-dependent snippets have synthetic URNs, original EN/AR text and individual UTF-8 text hashes; no real source was fetched or quoted.

## Response and scoring contract

Give the model only `prompt`, `inputs`, `allowed_tools` and `required_output`. Keep reference assertions, reference outputs and scoring metadata out of a scored model prompt. These examples and answers are public, so contamination is possible: scores on this sample cannot establish general capability or superiority.

The response must be exactly one JSON object with no Markdown fence or surrounding prose. Validate it against `required_output.schema` first. An invalid response fails all assertions; report schema validity separately. The scored output contains no unbounded additional narrative except the required English handoff draft.

For valid output, resolve each machine assertion’s `path` as a JSON Pointer (RFC 6901). Missing paths fail. Each assertion has weight 1:

1. `equals`: compare JSON values structurally. Object key order is irrelevant; array order matters; booleans never count as numbers.
2. `set_equals`: both values must be arrays; compare sets of canonical structural JSON values, ignoring item order. Object key order is irrelevant; numeric values such as `1` and `1.0` are equivalent; booleans remain distinct from numbers. `uniqueItems` in the output schema prevents duplicate credit.
3. `number_close`: both values must be finite JSON numbers, excluding booleans. Pass when `abs(actual - expected) <= absolute_tolerance`; no relative tolerance. JSON decimal literals are parsed without rounding and compared as exact rational values, avoiding binary-float and decimal-context rounding at tolerance boundaries.
4. `review`: award 1 only if every criterion is met using the supplied pass/fail anchors; otherwise 0. This occurs only for the English handoff text. Two independent bilingual reviewers should rate it blind; retain ratings and adjudicate disagreement. Until that review occurs, report deterministic assertion scores separately and label any single review provisional.

Report `passed_assertions`, `total_assertions`, `weighted_score = passed_assertion_weights / total_assertion_weights`, and `output_schema_valid` per task. `max_score` is the sum of weights, not a performance result. Report machine and review counts separately; omit a combined reviewed score while review is incomplete. Show EN and AR task results separately, plus pair-level agreement, if real runs are later recorded. Do not turn this score into an intelligence score or population-level finding.

The reference output is a checkable answer example, not a model result. The business case’s monthly calculation deliberately uses four weeks and treats saved minutes as reduced paid labor; it is a scenario estimate. The catalog flags the later duplicate SKU only and does not mutate data. The automation case plans a sequence and executes nothing.

## Optional offline scorer

Use Python 3.10 or later with the exact dependency `jsonschema==4.26.0`; this version was installed and verified for the implementation checks. Install it in an isolated environment before offline use:

```sh
python3 -m venv /tmp/jaai-scorer-venv
/tmp/jaai-scorer-venv/bin/python -m pip install jsonschema==4.26.0
```

Keep `scorer.py` beside the downloaded pack files. From that directory:

```sh
/tmp/jaai-scorer-venv/bin/python -B scorer.py --self-test
/tmp/jaai-scorer-venv/bin/python -B scorer.py \
  --task-id jaai-dec-001-en --output-file /absolute/path/to/answer.json
```

The dependency installation can require network access; scorer execution makes no network or model/API calls. The scorer reads local files, checks the dataset/schema hashes against the manifest, validates the pack and the supplied answer using JSON Schema Draft 2020-12 with exact numeric type handling, and prints JSON to standard output. Integral decimal values such as `9.0` satisfy `integer`; `9.0000000000000001` and booleans do not. It writes no files, changes no task data and never updates `results.json`. It refuses external schema references, duplicate JSON keys and nonfinite JSON numbers. Each JSON file is limited to 1 MiB and 64 nesting levels; each numeric literal to 128 characters, 100 mantissa digits and an exponent with absolute value at most 1,000. Parsing decimal literals as `Decimal` and comparing values and tolerances as exact `Fraction` values preserves all accepted literal digits without unbounded exponent expansion. An unknown task, malformed JSON or configuration error exits 2; a completed offline assessment exits 0 regardless of the answer’s score. A failing implementation self-test exits 1. `--self-test` refuses Python optimization (`-O` or `-OO`) because it requires its assertions to execute.

The report explicitly says `assessment_type: user_supplied_offline_assessment`, `recorded_model_run: false` and `independent_validation: false`. Machine numerators, denominators and the machine-only weighted score are separate from review assertions. Every `review` assertion stays `pending`, even when a draft appears correct or clearly wrong; `combined_reviewed_score` stays null. This is an assessment of a supplied answer, not a stored model run, independent review, performance finding or ranking. It cannot establish who or what produced the answer.

`--self-test` checks all 12 reference outputs and 84 machine assertions; leaves both paired language-review assertions pending; tests one concrete wrong answer in each of the six families; rejects four invalid output-schema controls and four invalid JSON controls; checks external-reference refusal and an independently calculated budget/savings case. Precision controls reject nonintegral `9.0000000000000001`, reject errors greater than `0.001` in both `4.0010000000000001` and a decimal with more than 28 digits, and accept the exact `4.001` boundary and integral `9.0`. Numeric-bound controls reject excessive digits and exponents. Reference checks and scorer self-tests are implementation checks, not benchmark executions or results. The public reference answers remain unsuitable as hidden evaluation data.

## Reproducible future runs

Freeze these files and record the SHA-256 of `tasks.json` before a run. Record dataset and scorer versions/hashes, harness commit/dependency lock, provider/model snapshot, exact prompt hash, generation settings, seed support, empty tool permissions, timestamps, output, failures, retries, latency and actual cost or a stated estimation method. Preserve every attempted run; do not discard failures. This pack includes no recorded model runs or cost data.

A future complete dataset, independent language review, run ledger and evidence report require separate versions. Use pair-level units for any within-pack uncertainty analysis; do not imply Jordan-wide representativeness, national leadership, clinical reliability, achieved AGI/SI or guaranteed employment.

## Authorship and terms

Author: Yousof Almalkawi, <https://steadywrk.app/team/yousof#person>. Sponsor: STEADYWRK. This is a self-sponsored artifact; authorship and sponsorship do not establish independent certification or measured capability. Independent bilingual review has not been performed.

Use is subject to <https://steadywrk.app/terms>. Provenance covers only these original synthetic examples and offline scorer code, imports no external copyrighted source text and grants no additional license to other website or repository material. The installed `jsonschema` dependency retains its own package license; it is not bundled here.

Initial version 0.1, authored 2 October 2026: six original paired examples and a truthful empty results ledger. No empirical findings published.

Protocol tooling addition, 2 October 2026: optional offline scorer with implementation self-tests. Original task data and the empty results ledger remain byte-for-byte unchanged.
