Trust Genome™ Methodology
The 12-dimension scoring framework APIR uses to compute every agent's trust score. Published, versioned, citable. Every score on the platform references a specific version of this document.
The 12 Dimensions
Identity
Confirms the agent is who it claims to be: verified operator org, platform registration, signing keys controlled by the registered party.
Behavioral
Stability of agent behaviour relative to its baseline. Drift detection, prompt-response consistency, side-effect patterns.
Security
Attack-surface posture: risk-tier classification, secret handling, OWASP Agentic Security top-10 controls.
Oversight
Human-in-the-loop coverage. Ghost Audit enabled, review frequency, escalation paths, override controls.
Compliance
Conformity with frameworks the operator opted into: EU AI Act, NIST AI RMF, ISO 42001, GDPR, HIPAA, SOC 2.
Fairness
Bias and equitable-outcome posture. APIR does not yet run independent bias testing, so this dimension is FIXED at 70 for every scored agent until it does. No agent can buy this score up.
Transparency
Model card published, data lineage documented, decision-logging coverage.
Accountability
Open-incident posture, escalation contacts current, audit-log coverage complete.
Data Integrity
Data classification declared, PII handling, lineage. Capped until integrity attestations are independently verified.
Performance
Operational reliability against committed thresholds.
Resilience
Monitoring active, failover behaviour, kill-switch wired.
Scoring
The Trust Genome is a documented derived composite, not 12 independent sensors. Each of the 12 dimensions is scored from the agent's own live evidence, up to that dimension's published cap. A dimension APIR cannot fully attest cannot pay out a perfect score: the caps are the honesty, not a limitation we hide.
Evidence gate: an agent is only scorable once it has produced at least 5 real behavioral events (actions, tasks, or decision-class Black Box records). Monitoring sweeps alone never make an agent scorable.
The overall Trust Score is the plain arithmetic mean of all 12 dimension scores: equal weighting, no hidden blend:
overall_trust_score = ( Σᵢ dimensionᵢ ) / 12
Because the caps sum to 1029, the maximum possible score is 85.75. That is deliberate: a perfect 100 would require attestations (independent bias testing, verified data-integrity chains, full mandate-level authorization proof) that the engine does not yet award. The ceiling rises only when the evidence model earns it, and every rise is a new version of this document.
Trust grade bands:
Note what the ceiling means for grade A: with a maximum of 85.75, grade A (≥85) is only reachable in a 0.75-point window that requires simultaneous perfection on every capped dimension: zero drift, zero anomalies, zero open incidents, compliant status, low risk tier, active monitoring, Ghost Audit on, full metadata, an issued passport and complete transparency fields. No agent in production has earned it yet, and we publish that rather than moving the band.
An agent only receives a grade once enough evidence has been collected to score it. Where the underlying signals are insufficient, the genome is marked insufficient_data and the agent renders as Not rated (grade NR), never a 0 or an F. Absence of evidence is reported as absence of evidence, not as failure.
Reproducibility
Every trust_genomes row in APIR's database carries the methodology_version that produced it. When this methodology changes, the version bumps; historical scores stay tied to the version under which they were computed.
A third party with the 12 dimension scores can independently confirm any overall trust score we publish: sum them, divide by 12, and check the result against the published caps and grade bands. No hidden weights, no proprietary blend.
Version history
Citing this document
APIR Intelligence. (2026). Trust Genome Methodology v…. https://apir.ai/methodology