525 decision models across 28 domains
Any tool can return a score. These models are built for decisions that get reviewed later, by someone who was not in the room, so every run has to carry its own reasoning, its own uncertainty, and a record an auditor can reopen.
No model returns a bare number. Each run ships an explainability section that ranks the drivers behind the result and shows how much each one moved it, in the direction it moved it. For the deterministic methods (weighted sum, TOPSIS, AHP, cost-benefit) that decomposition is derived from the criteria weights and the normalised performance matrix, and the output says so rather than dressing it up.
Where a run produced a trained model artifact, you get true SHAP values on demand: per-observation contributions against a base value, rendered as a waterfall and a force plot, with global feature importance across the population. Computed once and stored beside the run, so the explanation you review later is the same one the model gave at the time.
A single percentage is not an assessment. Every run returns the overall score together with the individual factors it was built from, each with its own score and a note explaining what that factor measured, plus the formula that combined them. If the inputs were thin, the run says which ones and how that fed through.
Confidence is capped at 0.95 by design. No model here claims certainty, and a run that has not been backtested against your own historical outcomes reports its validation status as model-estimated rather than calibrated. That cap is a deliberate honesty constraint, not a limitation of the method.
Each run is written to an immutable record that holds the exact inputs it was given, the complete output it returned, the plugin and version that produced it, the decision structure it reasoned over, and the timestamps around execution. Inputs are echoed back verbatim rather than summarised, so a reviewer can reproduce the run instead of trusting a description of it.
Runs attach to the decision they informed, so the lineage runs both ways: from a decision to the evidence behind it, and from a model run to everything it went on to justify. Every run is also scanned for personal data before it is stored, and the classification travels with the record.
This is not a per-model variation. Whichever model you run, the result comes back in the same shape, so a reviewer learns to read one format instead of 28 domain conventions.
A model library is only worth what its weakest entry is worth. Every model passes the same six stages before release, and stays under observation afterwards.
Every model declares a typed input schema and output schema before any logic is written. Inputs are validated against those constraints at run time, so a malformed run fails loudly at the boundary instead of quietly producing a number.
Models are exercised in three rounds against canonical datasets: correctness on known inputs, behaviour at the boundaries, and stability across the realistic range. A model that only works on the happy path does not ship.
An independent AI evaluator grades each model from A to F on methodology soundness, blind spots, and whether its stated confidence is honestly calibrated. The rubric is fixed. Models are improved to earn the grade, never the reverse.
Released models are cryptographically signed and version-pinned. The runtime verifies the signature before execution and records which exact version ran, so a result can always be traced to the code that produced it.
Execution happens in a sandbox with enforced time and resource limits. One model cannot reach another tenant's data, and a runaway run is contained rather than shared.
Live models are watched for drift and bias after release. Prediction accuracy is tracked against real outcomes, and calibration reports show where the model has been over-confident or under-confident in practice.
Bringing your own model does not opt you out of any of this. External and custom models register against the same schema contract, return the same twelve sections, and carry the same run record, so a model you trained is reviewed on the same terms as one of ours.