Validation System and fidelity requirements
Defines fidelity cards, qualification, analytical benchmarks, user-authored regression tests, linearization, identifiability, optional independent evidence comparisons, uncertainty, calibration evidence, and convergence expectations.
Section relationships
| ID | Name | Priority | Requirement | Verification |
|---|---|---|---|---|
| VAL-001 | Fidelity card | MVP | Every completed run SHALL include a fidelity card listing active models and effects, solver settings, engineering-property transformations and final dispositions, inactive engineering properties, inactive or unsupported effects, approximations, parameter evidence, validity warnings, and applicable evidence references or their explicit absence. | Inspection |
| VAL-002 | Per-capability qualification | MVP | Qualification status SHALL be assigned to a model capability and use case rather than to BeforeMetal as one undifferentiated product. | Inspection |
| VAL-003 | Qualification vocabulary | MVP | BeforeMetal SHALL define and use documented statuses for unqualified, analytically or numerically verified, reference-benchmarked, calibrated, compared with independent reference data, and physically validated models where applicable. | Inspection |
| VAL-004 | Observable-specific comparison | MVP | Any validation or independent-reference comparison SHALL report error or a pass criterion for named observables, operating conditions, and identified evidence sources rather than one universal accuracy percentage. | Analysis |
| VAL-005 | Analytical benchmark suite | MVP | The Validation System SHALL include the versioned analytical cases and tolerances approved under TBD-VAL-ANALYTICAL for free fall, projectile motion, pendulum motion, energy behavior, and supported actuator equations. | Analysis |
| VAL-006 | Contact benchmark suite | MVP | The Validation System SHALL include the versioned contact cases and tolerances approved under TBD-VAL-CONTACT for restitution, incline stick-slip, stacking, penetration, friction behavior, and repeatability. | Analysis |
| VAL-007 | Joint benchmark suite | MVP | The Validation System SHALL include the versioned joint cases and tolerances approved under TBD-VAL-JOINT for kinematics, limits, drives, constraint drift, and articulated chains. | Analysis |
| VAL-008 | Device benchmark suite | MVP | The Validation System SHALL include the versioned device cases and tolerances approved under TBD-VAL-DEVICE for motors, transmissions, batteries, power balance, saturation, and latency. | Analysis |
| VAL-009 | Sensor benchmark suite | MVP | The Validation System SHALL include the versioned sensor cases and tolerances approved under TBD-VAL-SENSOR for timing, frames, bias, drift, noise, quantization, saturation, latency, and dropout. | Analysis |
| VAL-010 | Weather benchmark suite | MVP | The Validation System SHALL include the versioned cases and tolerances approved under TBD-VAL-WEATHER for fields, wind load, wet friction, and the reference range-sensor weather effect. | Analysis |
| VAL-011 | Timestep convergence | MVP | Each released mechanics profile SHALL satisfy the observable, timestep range, procedure, and convergence criteria approved under TBD-VAL-CONVERGENCE. | Analysis |
| VAL-012 | General solver-parameter sensitivity | Beta | Each released engineering solver profile beyond the MVP reference configurations covered by ACCEPT-009 SHALL document sensitivity to relevant tolerance and iteration settings. | Analysis |
| VAL-013 | Mesh convergence | Future | Mesh-based structural, fluid, or field models SHALL include documented spatial convergence evidence before being marked validated. | Analysis |
| VAL-014 | Cross-backend comparison | Beta | Equivalent scenarios SHALL be runnable across compatible backends and compared by named observables and tolerances. | Analysis |
| VAL-015 | Cross-backend is not truth | MVP | User-facing reports SHALL state that agreement between simulators is not physical validation. | Inspection |
| VAL-016 | Calibration dataset | MVP | A calibrated model SHALL identify the immutable observations used for parameter fitting. | Inspection |
| VAL-017 | Independent evaluation dataset | MVP | A model claimed as physically validated or compared with independent reference data SHALL identify, with provenance, the evaluation observations or reference cases not used during calibration, tuning, or acceptance-rule selection. | Inspection |
| VAL-018 | Calibration-evaluation separation | MVP | When the reference workflow uses calibration and evaluation datasets, it SHALL prevent designated evaluation observations from contributing to fitting, tuning, or acceptance-rule selection unless the dataset designation is explicitly revised before that use. | Test |
| VAL-019 | Uncertainty propagation | Beta | A validation report SHALL account for declared parameter and measurement uncertainty when comparing predicted and observed quantities. | Analysis |
| VAL-020 | Confidence reporting | Beta | Stochastic result summaries SHALL state sample count, interval method, confidence or credibility level, and failed-sample handling. | Analysis |
| VAL-021 | Validity-envelope enforcement | MVP | A run outside a model's declared validity envelope SHALL emit a named warning or compilation error according to policy. | Test |
| VAL-022 | Validation provenance | MVP | A validation or independent-reference result SHALL record model revision, software revision, settings, compute hardware where applicable, evidence-source identity, publisher or owner, dataset revision, license or access terms, procedure, and author. | Inspection |
| VAL-023 | Validation regression gate | MVP | A change that moves a committed benchmark outside its approved envelope SHALL block release unless the envelope and rationale are versioned. | Test |
| VAL-024 | Independent decision-ranking validation | Beta | A release claiming externally validated design ranking SHALL test whether BeforeMetal ranks selected alternatives consistently with a versioned independent comparative dataset that was not used to select, tune, or threshold the ranking rule. | Analysis |
| VAL-025 | Reference evidence sources | MVP | For every MVP qualification claim, the product definition SHALL identify an applicable versioned analytical, published, manufacturer, community, or independently collected evidence source and its provenance or explicitly declare the affected capability unqualified. | Inspection |
| VAL-026 | Invalid numeric state | MVP | A run SHALL detect and terminate with diagnostics on required NaN, infinity, non-physical mass state, or unrecoverable solver failure. | Test |
| VAL-027 | Energy and constraint budget | Beta | A mechanics validation report SHOULD include documented energy drift, constraint error, penetration, and impulse or force budgets where meaningful. | Analysis |
| VAL-028 | Community model status | Beta | Community-contributed models SHALL display their evidence status and SHALL not inherit built-in validation claims automatically. | Inspection |
| VAL-029 | External solver evidence | External | An external solver adapter SHALL expose enough identity and configuration data to associate its results with applicable vendor or BeforeMetal validation evidence. | Inspection |
| VAL-030 | Fidelity comparison | Beta | A user SHALL be able to compare Preview, Engineering, Validation, or extension-defined fidelity configurations without assuming their results are interchangeable. | Demonstration |
| VAL-031 | Fidelity conformance status | MVP | A run that overrides a selected profile's minimum qualification, convergence, uncertainty, or validity policy SHALL be marked nonconforming and SHALL not be presented as satisfying that profile. | Test |
| VAL-032 | Evidence source classification | MVP | Every qualification or validation reference SHALL record both the claim class it can support and its evidence source class: analytical derivation, standard benchmark, published experiment, manufacturer or vendor source, community dataset, independent reference result, cross-solver comparison, or project-specific measurement. | Inspection |
| VAL-033 | Declarative simulation regression case | Beta | A user-authored simulation regression case SHALL reference an experiment revision, execution mode, seed or declared seed-and-repetition plan, optional runtime-action plan, stop conditions, and declarative assertions over termination, diagnostics, events, metrics, committed state, or traces with explicit units, frames, comparators, tolerances, time windows, and statistical criteria where applicable. | Test |
| VAL-034 | Headless regression-suite execution | Beta | BeforeMetal SHALL execute a finite user-authored regression suite without a UI, create a fresh world instance for each case, use stable case selection and ordering, fail the suite when any required case fails or remains incomplete, and emit a machine-readable suite summary. | Test |
| VAL-035 | Regression baseline versioning | Beta | Every expected regression baseline SHALL have a content identity, and changing an accepted baseline SHALL create a new version with explicit approval and rationale while preserving prior results and evidence. | Test |
| VAL-036 | Dynamic-analysis qualification | Beta | Every released General Robotics Beta dynamic-analysis profile SHALL pass the analytical or reference cases and criteria approved under TBD-VAL-DYNAMIC-ANALYSIS for generalized mass, bias, gravity, force mapping, inverse-dynamics consistency, derivatives, local linear prediction, and practical-identifiability diagnostics. | Analysis |
Change rationale (VAL-001, VAL-003, VAL-004, VAL-012, VAL-017, VAL-018, VAL-022, VAL-024, VAL-025, VAL-032): ADR-0001 replaces mandatory team-run physical qualification with explicit evidence classes and honest unqualified states. Analytical and numerical qualification remain mandatory; published, manufacturer, community, independent, and optional physical data may strengthen named claims. Physically validated status still requires physical evidence, and cross-solver agreement remains comparison evidence rather than physical truth. External decision-ranking validation moves to Beta because it is conditional on a suitable independent comparative dataset. The completed fidelity card discloses property closure and evidence absence, while general solver-profile sensitivity remains Beta beyond the reference-case convergence required for MVP.
Change rationale (VAL-033–VAL-036): BeforeMetal's own release benchmarks do not replace user-defined robot and controller regression tests. Declarative assertions, explicit stochastic criteria, immutable expected baselines, and machine-readable suite outcomes make those tests reproducible without arbitrary embedded scripts or an assumption of bitwise equality. Passing a regression establishes behavioral conformance, not physical validation. Dynamic-analysis qualification separately bounds derivatives, linear models, and identifiability to approved local cases rather than claiming validity across contact changes or outside a declared operating region.
Generated from the canonical specification. Edit section metadata or prose in docs/requirements.md; the website rebuilds this page and its relationships automatically.