The operating problem is frozen before a model is selected.
CAPITAL financial decision systems.
CAPITAL evaluates whether models separate observed market or accounting data from assumptions, causal hypotheses, model output, valuation conclusions, and decision authority.
Finance punishes false precision. The benchmark rewards disciplined mechanism identification and controlled decisions under non-stationarity.
Inputs, hidden facts, allowed records, and prohibited data are bounded.
Permitted tools and action surfaces are stated before the run.
Result fields, confidence, timing, and reproduction state are recorded.
Unsafe action, false evidence, or authority violation ends the run.
The packet separates model behavior from harness, tool, and evaluator state.
Benchmark first. Model second. Evidence always attached.
AMERICA-SIM and the CognitivityX benchmark families exist to prevent result laundering. The same case, evidence, tools, timing, scoring, and hard fails apply before the run begins.
The operating problem is frozen before a model is selected.
Inputs, hidden facts, allowed records, and prohibited data are bounded.
Permitted tools and action surfaces are stated before the run.
Result fields, confidence, timing, and reproduction state are recorded.
Unsafe action, false evidence, or authority violation ends the run.
The packet separates model behavior from harness, tool, and evaluator state.
Benchmark packet
Each benchmark page states how the mission is locked, scored, failed, and reported.
- Task family
- CAPITAL financial decision systems
- Example case
- Valuation or risk case separating observed data, assumptions, model output, decision authority, and regulatory boundary.
- Scoring
- Correctness, evidence discipline, constraint satisfaction, calibration, action authority, and hard-fail avoidance.
- Hard fail
- Unsupported authority, hidden constraint violation, fabricated evidence, unsafe release, or failure to escalate.
- Report
- Inputs, prompt packet, evaluator, score, failure labels, confidence interval, and reproduction notes.
Rigor has to be designed into the mission surface.
Mission definition
CAPITAL evaluates whether models separate observed market or accounting data from assumptions, causal hypotheses, model output, valuation conclusions, and decision authority. The use case determines the operating contract before the model is selected.
Comparator discipline
Each model faces the same case, available information, and acceptance conditions. Differences in internal architecture, reasoning efficiency, or ability to use the supplied information are part of what the benchmark measures.
Result discipline
Controlled internal results, reproduced runs, and external vendor benchmarks are kept in separate evidence classes. A result is never upgraded by presentation alone.
Hard fails are part of the intelligence test.
A result can look fluent and still fail the mission. The benchmark family makes disallowed action, broken authority, unsupported evidence, and unsafe release visible as first-class outcomes.
The system acts outside the mission boundary or role permission.
A claim is presented as supported when the record does not support it.
The result ignores a fixed operational, legal, physical, or statistical limit.
The model releases action when uncertainty or policy requires deferral.
Benchmarks test the model. They do not become the model.
CognitivityX keeps Altheon, Berni, Aletheion, STRATA, STRATOS, and benchmark evidence separate so a visitor can see what is being measured, what is being controlled, and what is being released.
Example case
Valuation or risk case separating observed data, assumptions, model output, decision authority, and regulatory boundary.
Result status
Scores are separated by controlled run, reproduced run, reference evidence, and production verification.
Hard-fail discipline
The page separates task, inputs, scoring, hard fail, result class, and reproduction notes.
Contact us for access
Use the access route for model review, benchmark packet, system walkthrough, or institutional collaboration.
Inspect evidence
- Status
- Protocol visible
- Primary artifact
- CAPITAL financial decision systems.
- Evidence class
- Task family
- Source
- Benchmarks
- Date
- 2026-09-22
- Related paper
- AMERICA-SIM Standard
- Benchmark
- CAPITAL
- Known limitation
- Public pages exclude protected corpora, hidden cases, private scoring weights, partner data, and credentials.
- Required next proof
- Inputs, prompt packet, evaluator, score, failure labels, confidence interval, and reproduction notes.
Request benchmark access.
For benchmark packets, methodology review, matched comparator runs, or private evaluation discussion, send the mission context, data boundary, scoring need, and review timeline.