Model Library

    525 decision models across 28 domains

    All (525)
    AI & Agentic Governance (32)
    Banking & Credit (12)
    Benefits & Total Rewards (51)
    Customer & Revenue (10)
    Data Governance (15)
    Decision Governance (14)
    Energy & Industrial (10)
    ESG & Climate (11)
    Finance & FP&A (43)
    Financial Risk & Treasury (9)
    Governance & Risk (23)
    Grant Compliance (12)
    Health & Safety (24)
    HR & People (69)
    Insurance & Underwriting (13)
    Legal & IP (8)
    M&A (15)
    Operations (24)
    Physical Autonomy & Robotics (6)
    Procurement & Operations (10)
    Product & Innovation (9)
    Provider & Clinical Ops (13)
    Public Sector & Education (13)
    Real Estate & Capital Projects (12)
    Risk, Security & Resilience (14)
    Sales & Revenue (19)
    Strategy (24)
    Strategy & Board (10)
    525 Models·14 Methods·28 Domains

    Behind every model

    A number is not an answer

    Any tool can return a score. These models are built for decisions that get reviewed later, by someone who was not in the room, so every run has to carry its own reasoning, its own uncertainty, and a record an auditor can reopen.

    Explainability

    Every run decomposes its own answer

    No model returns a bare number. Each run ships an explainability section that ranks the drivers behind the result and shows how much each one moved it, in the direction it moved it. For the deterministic methods (weighted sum, TOPSIS, AHP, cost-benefit) that decomposition is derived from the criteria weights and the normalised performance matrix, and the output says so rather than dressing it up.

    Where a run produced a trained model artifact, you get true SHAP values on demand: per-observation contributions against a base value, rendered as a waterfall and a force plot, with global feature importance across the population. Computed once and stored beside the run, so the explanation you review later is the same one the model gave at the time.

    • Ranked driver decomposition on every run
    • SHAP waterfall and force plot for trained models
    • Global feature importance across observations
    • States when a value is weight-derived, not SHAP
    Confidence

    A confidence score you can take apart

    A single percentage is not an assessment. Every run returns the overall score together with the individual factors it was built from, each with its own score and a note explaining what that factor measured, plus the formula that combined them. If the inputs were thin, the run says which ones and how that fed through.

    Confidence is capped at 0.95 by design. No model here claims certainty, and a run that has not been backtested against your own historical outcomes reports its validation status as model-estimated rather than calibrated. That cap is a deliberate honesty constraint, not a limitation of the method.

    • Per-factor breakdown, each with score and rationale
    • The explicit formula that produced the overall score
    • Data-quality warnings tied to specific inputs
    • Validation status: model-estimated or backtested
    Audit trail

    Inputs and outputs, kept whole

    Each run is written to an immutable record that holds the exact inputs it was given, the complete output it returned, the plugin and version that produced it, the decision structure it reasoned over, and the timestamps around execution. Inputs are echoed back verbatim rather than summarised, so a reviewer can reproduce the run instead of trusting a description of it.

    Runs attach to the decision they informed, so the lineage runs both ways: from a decision to the evidence behind it, and from a model run to everything it went on to justify. Every run is also scanned for personal data before it is stored, and the classification travels with the record.

    • Verbatim input echo, not a summary
    • Full output payload retained with the run
    • Plugin identity, version and execution timestamps
    • PII classification attached to every record
    • Two-way linkage between runs and decisions
    The output contract

    Twelve sections, on every one of the 525 models

    This is not a per-model variation. Whichever model you run, the result comes back in the same shape, so a reviewer learns to read one format instead of 28 domain conventions.

    1. 01Executive summaryHeadline, risk level, key metrics, action items
    2. 02Key findingsRanked, each with severity and its supporting evidence
    3. 03Confidence scoringOverall score, factor breakdown, warnings
    4. 04Policy evaluationThreshold checks with pass or fail status
    5. 05RecommendationsPrioritised, each with rationale and expected impact
    6. 06VisualizationsChart specifications the UI renders directly
    7. 07Assumptions and limitationsWhat the model took as given, and where it stops
    8. 08Methodology documentationWhy this method, and what was rejected
    9. 09Legal disclaimerGovernance and professional-advice notice
    10. 10Data quality assessmentCompleteness, provenance, known issues
    11. 11Sensitivity summaryWhich parameters the result actually turns on
    12. 12ExplainabilityFeature importance and driver decomposition
    How they are built and maintained

    Nothing reaches this catalog ungraded

    A model library is only worth what its weakest entry is worth. Every model passes the same six stages before release, and stays under observation afterwards.

    Specify

    Every model declares a typed input schema and output schema before any logic is written. Inputs are validated against those constraints at run time, so a malformed run fails loudly at the boundary instead of quietly producing a number.

    Test

    Models are exercised in three rounds against canonical datasets: correctness on known inputs, behaviour at the boundaries, and stability across the realistic range. A model that only works on the happy path does not ship.

    Grade

    An independent AI evaluator grades each model from A to F on methodology soundness, blind spots, and whether its stated confidence is honestly calibrated. The rubric is fixed. Models are improved to earn the grade, never the reverse.

    Sign

    Released models are cryptographically signed and version-pinned. The runtime verifies the signature before execution and records which exact version ran, so a result can always be traced to the code that produced it.

    Isolate

    Execution happens in a sandbox with enforced time and resource limits. One model cannot reach another tenant's data, and a runaway run is contained rather than shared.

    Monitor

    Live models are watched for drift and bias after release. Prediction accuracy is tracked against real outcomes, and calibration reports show where the model has been over-confident or under-confident in practice.

    Bringing your own model does not opt you out of any of this. External and custom models register against the same schema contract, return the same twelve sections, and carry the same run record, so a model you trained is reviewed on the same terms as one of ours.

    Ready to run your first model?

    See how DecisionLedger AI transforms your decision-making.