Skip to content

FilingsCheck · 2026-07-filingscheck-v1.0.1-deployed

Every grade on this page is provable.

FilingsCheck asks frontier models for figures reported in SEC filings and verifies every claim against the primary source as filed. Each grade is sealed in a witnessed, append-only audit log. This edition ran against the deployed production node. Nothing below asks for trust.

  • Published 2026-07-28
  • 485 sealed grades
  • $0.08118625 total elicitation spend

Engine scorecard

Read this first

Adversarial probes measure Artizar’s own accuracy, not the models under test. Every run injects known-answer probes (true controls plus deliberate falsehoods) and grades the engine on catching them.

adversarial probes · all runssealed
372 / 372
probes graded as expected, across all runs
0
false greens: planted falsehoods graded confirmed
0
false alarms: true controls graded contradicted
  • claude-haiku-4-593 / 93 · 0 false greens · 0 false alarms
  • claude-sonnet-593 / 93 · 0 false greens · 0 false alarms
  • gpt-5.4-mini93 / 93 · 0 false greens · 0 false alarms
  • gpt-5.493 / 93 · 0 false greens · 0 false alarms

A false green is a planted falsehood graded confirmed. It is the failure this system exists to not have.

Model matrix

Four models, one standard of proof

FilingsCheck asks each model for exactly the fact the engine verifies; the model may abstain. The headline is supported over asserted: abstentions are honest and leave the denominator. A low headline is not an engine failure. It is the engine catching a model’s errors, with a sealed record for each catch.

claude-haiku-4-5-20251001anthropic
supported
7
contradicted
23
abstained
0
headline
23.3% (7/30)
cost
$0.011168
claude-sonnet-5anthropic
supported
22
contradicted
6
abstained
2
headline
78.6% (22/28)
cost
$0.038463
gpt-5.4-mini-2026-03-17openai
supported
15
contradicted
14
abstained
1
headline
51.7% (15/29)
cost
$0.00770775
gpt-5.4-2026-03-05openai
supported
24
contradicted
2
abstained
4
headline
92.3% (24/26)
cost
$0.0238475

Reproduction

Verify this edition yourself

Each run ships as a sealed bundle. One command per bundle checks, offline, every grade’s full chain: the Ed25519 seal signature, the commitment recompute, the audit-log entry hash, the witnessed inclusion proof, report-to-verdict consistency, and the prompt-version label against the claims. Any failure is a named error.

fcbundle/v1 · offline verificationone command per run
python -m artizar verify-bundle bundles/anthropic-claude-haiku-4-5
python -m artizar verify-bundle bundles/anthropic-claude-sonnet-5
python -m artizar verify-bundle bundles/openai-gpt-5.4-mini
python -m artizar verify-bundle bundles/openai-gpt-5.4
key_provenance
deployed (operator-published; https://artizar.ai/.well-known/artizar-keys.json; S74 ceremony 2026-07)

The verification keys are public at the address above. The sealed bundles themselves are shared with design partners. Request access.

Version labels

What, exactly, was measured

Comparisons across editions are apples to oranges unless these labels match. Every label is read from the sealed artifacts, never typed by hand, and a mixed-version edition refuses to build.

2026-07-filingscheck-v1.0.1-deployedread from sealed artifacts
taskset
taskset-v6
prompt
v2 · template sha256 fb0029a26d6e91aea980639ce2e4b3b8cb2f2ed7c663d22bcc6f4875a9dfa6a7
scoring_policy
policy-v7
period_binding
binding-v3
metric_lexicon
lexicon-v1
harness
filingscheck/v2
verdict_schema
verdict-v3
pricing_table
prices-2026-06-24
bundle_format
fcbundle/v1

Reading the results

How to read FilingsCheck

  • Method

    Models are asked for the fact itself.

    Each task elicits exactly the figure the engine verifies: a reported value from a specific SEC filing. The model answers or abstains; the engine checks the answer against the filing as filed and seals the grade.

  • Headline

    Supported over asserted.

    Abstentions are honest and leave the denominator. The headline measures what a model chose to assert, not what it declined to guess.

  • Reading it

    A low headline is the engine working.

    When a model asserts a wrong figure, the engine contradicts it and the record shows the catch. The scorecard above, with zero false greens across every run, is what makes the matrix worth reading.

Benchmark your model against the record.

Artizar is pre-launch and onboarding a small set of early partners.

Request early access