How the Compassion Benchmark works
Most institutional ratings measure what organizations say. The Compassion Benchmark measures what they do when compassion is costly — scored on a 0–100 scale built from 8 dimensions and 40 evidence-checked subdimensions.
What every published score guarantees
- Evidence-grounded — every score traces to documented public evidence across a 5-tier hierarchy, not opinion or self-report.
- Adversarially tested — performance only counts when it held up under cost or pressure (the pressure-test principle).
- Reproducible — the same 8 dimensions and 40 subdimensions are applied to every entity, so scores are comparable across sectors.
- Independent — no entity pays to be included, scored higher, or have findings withheld.
- Contestable — methodology version, evidence tiers, and scoring are published so any score can be checked and challenged.
If you have 3 minutes — the short version
Every published score carries five guarantees. Here is what each one means in practice:
- Evidence-grounded. Every score traces to documented public evidence across a 5-tier evidence hierarchy — not opinion, self-report, or press releases. Adverse events and positive structural evidence are both actively searched.
- Adversarially tested. High scores only count if the compassionate behavior held under cost or pressure. See the pressure-test principle.
- Reproducible. The same 8 dimensions and 40 subdimensions apply to every entity. Scores can be compared across sectors because the framework does not change by sector.
- Independent. No entity pays to be included, scored higher, or have findings withheld. The independence policy is the load-bearing trust commitment.
- Contestable. Methodology version, evidence tiers, and the scoring formula are all published. Any score can be checked and challenged.
Core methodological principle
The pressure-test principle
Every dimension is assessed under adversarial conditions. For each subdimension, assessors look for at least one documented case where compassionate behavior was costly, legally risky, or institutionally inconvenient. If no such case exists, the maximum subdimension score is capped at Developing, even when the entity appears strong under favorable conditions.
In plain terms: high performance when it is easy is not treated as sufficient evidence of compassionate institutional character.
See this rule applied: Hugging Face maintains an Exemplary designation across 7 of 8 dimensions, with pressure-test evidence present for each.
How scores are built
From individual subdimension anchors to the 0–100 composite.
The scoring pipeline
40 subdimensions (0–5) → 8 dimensions (/10) → base composite 0–100 via ((avg−1)/4×100) → +premium /10 → composite /100 → band
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — The scoring pipeline. compassionbenchmark.com/methodology. CC-BY 4.0.
Common scoring model
Each subdimension is scored on a 0–5 anchored behavioral scale. The five subdimensions within a dimension are summed and converted into a dimension score out of 10. The eight dimension scores are averaged on the 0–5 scale and converted to a base composite via ((average − 1) ÷ 4) × 100 (a 0–100 value); an integration premium of up to 10 points is then added and the result clamped to a 0–100 maximum.
((avg − 1) ÷ 4) × 100 = base composite (0–100) + integration premium (0–10) = composite 0–100
A score of 0 represents active documented harm and requires lead assessor co-sign. A score of 4 or 5 is provisional unless there is pressure-test evidence.
Integration premium logic
The bonus is 10 × a consistency factor (lower variance across dimensions scores higher) × a balance factor (fewer weak dimensions scores higher), so a balanced 70/70 profile can beat a spiky 90/40 one; a single dimension at zero sets the bonus to 0.
Consistency is rewarded: strong, even performance across all eight dimensions earns up to +10 points; any dimension at zero (active harm) cancels the bonus.
Standard deviation → consistency factor
Lower variance across 8 dimension scores earns a higher premium multiplier.
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — Standard deviation → consistency factor. compassionbenchmark.com/methodology. CC-BY 4.0.
See this rule applied: Costco earns a reduced integration premium because its EQU and SYS dimension scores are below the 4.0 balance threshold.
A real scored entity: Abridge
Walking the full scoring chain — anchors → dimension → base → integration premium → composite → band — using a real entity from the AI Labs index.
Step 1 — Subdimension anchors → dimension scores
Each of the 8 dimensions has 5 subdimensions, each scored 0–5 against behavioral anchors. The 5 subdimension scores sum to a dimension score out of 25, then convert to /10 (÷2.5). Abridge's assessed subdimension averages:
| Dimension | Subdim avg (0–5) | Dim score (/10) | Anchor level |
|---|---|---|---|
| AWRAwareness | 3.5 | 7.0 | Developing → Established |
| EMPEmpathy | 3.5 | 7.0 | Developing → Established |
| ACTAction | 3.5 | 7.0 | Developing → Established |
| EQUEquity | 3.0 | 6.0 | Developing |
| BNDBoundaries | 3.5 | 7.0 | Developing → Established |
| ACCAccountability | 3.5 | 7.0 | Developing → Established |
| SYSSystemic Thinking | 3.5 | 7.0 | Developing → Established |
| INTIntegrity | 3.5 | 7.0 | Developing → Established |
Step 2 — Base composite
Average the 8 dimension scores on the 0–5 scale: (3.5 + 3.5 + 3.5 + 3.0 + 3.5 + 3.5 + 3.5 + 3.5) ÷ 8 = 3.44. Base composite = ((3.44 − 1) ÷ 4) × 100 = 60.9.
Step 3 — Integration premium
The formula is: premium = 10 × consistencyFactor × balanceFactor (with any single dimension at exactly 0 zeroing the premium entirely).
Consistency factor — standard deviation across the 8 subdimension averages [3.5, 3.5, 3.5, 3.0, 3.5, 3.5, 3.5, 3.5] ≈ 0.17 (σ ≤ 1.5 → factor = 1.0).
Balance factor — counts dimensions below the 4.0 threshold (on the 0–5 scale). All 8 of Abridge's subdimension averages are below 4.0 (highest is 3.5) → 8 weak dimensions → factor = max(0, 1 − 8 × 0.2) = max(0, −0.6) = 0.0.
Formula substitution: 10 × 1.0 × 0.0 = 0.0.
Teaching note:Abridge earns no integration premium because the premium rewards entities whose dimension scores clear the 4.0-of-5 bar — a threshold that signals balanced, cross-dimensional strength. All of Abridge's dimensions sit at 3.0–3.5, so the balance factor collapses to zero and the premium is 0.0. The premium is deliberately hard to earn; see the three-profile diagram above for examples of profiles that do earn a positive premium.
Step 4 — Composite and band
60.9 (base composite) + 0.0 (premium) = 60.9 composite (clamped to a 0–100 maximum).
60.9 falls in the Established band (60–80): practices are systematic, documented, and supported by consistent evidence — but the sub-4.0 average on Equity (EQU 3.0) signals a dimension requiring further improvement.
Rubric anchors and score bands
The Human Assessment Battery uses universal behavioral anchors at the subdimension level and a five-band public interpretation model at the composite level.
What each score level means
The 0–5 behavioral anchor scale. Every subdimension is scored against these six anchors.
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — What each score level means. compassionbenchmark.com/methodology. CC-BY 4.0.
What the composite score means
The five composite bands mapped across the 0–100 scale.
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — What the composite score means. compassionbenchmark.com/methodology. CC-BY 4.0.
These labels describe two different things.Anchor levels score a single subdimension (0–5); bands describe the whole-entity composite (0–100). They share some names by design but are not the same measurement — an entity can have several ‘Established’ subdimensions and still land in a lower band.
| Score | Anchor level | Meaning |
|---|---|---|
| 0 | 0 ·Active Harm | Specific documented harm in the domain; lead assessor co-sign required. |
| 1 | 1 ·Absent | No meaningful capacity exists. |
| 2 | 2 ·Minimal | Nominal capacity exists but fails under pressure and does not produce consistent real-world outcomes. |
| 3 | 3 ·Developing | Good-faith capacity exists in some cases, but not consistently or comprehensively. |
| 4 | 4 ·Established | Consistent operational capacity across most cases; community confirms positive experience. |
| 5 | 5 ·Exemplary | Outstanding independently verified performance sustained under significant pressure. |
Evidence hierarchy
The benchmark deliberately differentiates evidence by independence and reliability. Strong scores require stronger evidence.
5-tier evidence trust pyramid
Tiers narrower at top = higher weight. Evidence beats aspiration: where claims diverge from lived experience, the methodology scores the world as encountered.
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — 5-tier evidence trust pyramid. compassionbenchmark.com/methodology. CC-BY 4.0.
Recency and evidence decay — what the 14-day window governs
The 14-day window is the scan cadence, not the lifespan of the evidence base. Each nightly cycle looks for material new evidence within the most recent 14 days; that window governs what can trigger a re-assessment, not what a score rests on.
- Baseline scores draw on a multi-year evidence base. Assessments routinely cite findings from prior years as the load-bearing evidence for a dimension. Harvard's Empathy and Accountability scores rest on standing findings (an OCR Title VI/IX violation, the Comaroff harassment case) that predate the current window; because no new conduct toward students emerged, those dimensions held.
- Adjudicated findings persist. A court ruling, regulatory finding, or settlement remains part of the evidence base until superseded by documented change — it does not expire on a 14-day clock.
- Uncorroborated allegations decay. Evidence that is not corroborated and not adjudicated carries less weight over time and does not accumulate into a score change on its own. This is the same discipline that keeps pre-adjudication probes as evidence-tier upgrades rather than composite movements until adjudication occurs.
Served population — whose experience the evidence search is scoped to
The evidence search is scoped to how the entity treats its served population (the subjects defined in the attribution rule):
| Index | Population the evidence search covers |
|---|---|
| Countries | Residents and people under the state's effective control or authority |
| Cities (US & global) | Residents and people the city serves |
| Companies (Fortune 500) | Workers, customers, and the communities the company operates in |
| AI labs | Users and the broader society affected by the lab's systems |
| Robotics labs | Users and the broader society affected by the lab's systems |
| Universities | Students, workers (faculty and staff), and surrounding communities |
Positive-evidence search — countering the desk-research downward bias
Desk-based research has a built-in downward bias: adverse events (lawsuits, probes, strikes, casualties) are reported far more aggressively than the quiet existence of working structures. To counter this, assessors must actively search for structural positive evidence, not only adverse news.
Positive evidence is structural and verifiable — published outcome data acted upon, independent audits with corrective action, durable access and equity infrastructure, worker-voice gains, and commitments held at real cost — not press releases or mission statements.
Without an active positive search, the benchmark would systematically under-score entities that are quietly competent and over-weight whoever happened to be in the news. University labor wins (graduate-worker contract gains, faculty unionization) and Harvard's choice to bear cost rather than settle a compliance demand are both examples of positive signals that require deliberate search to surface.
Who is scored & whose harm counts
Two questions must be settled before any harm event can be scored: who is the subject, and which actor does a given harm belong to. The rules below are testable and applied uniformly across indexes.
Scored subject by index type
The subject is the population the entity is responsible for recognizing, responding to, and not depleting. Scoring asks how the entity treats that population.
| Index | Scored subject (the population the entity is responsible for) |
|---|---|
| Countries | Residents and people under the state's effective control or authority |
| Cities (US & global) | Residents and people the city serves |
| Companies (Fortune 500) | Workers, customers, and the communities the company operates in |
| AI labs | Users and the broader society affected by the lab's systems |
| Robotics labs | Users and the broader society affected by the lab's systems |
| Universities | Students, workers (faculty and staff), and surrounding communities |
The victim / perpetrator test
For any harm event in the evidence window, decide where the harm belongs before scoring it:
- Identify the actor who caused the harm. Ask who did the harmful thing, not merely where the harm landed.
- If the entity being scored is the actor — the harm reflects the entity's own conduct toward its subject population — the event is scored against the entity (subject to the evidence and adjudication standards).
- If an external actor is the cause and the scored entity is the victim — the event is attributed to the perpetrator, not the entity, and is not scored as a compassion failure of the entity.
- Apply the rule symmetrically. The reverse also holds: if an entity inflicts the same category of harm on its own people by its own choice (rather than having it imposed from outside), that conduct is scored against the entity.
In the 2026-06-20 cycle this rule was applied at index scale: the entire June 2026 federal campaign against US universities (DOJ funding suits, grant freezes, admissions compliance reviews) was treated as external action inflicted onthe universities and attributed to the federal government — so none of it lowered a university's score. The same convention attributes strikes on Ukraine to Russia, leaving Ukraine's own conduct profile unchanged.
Worked example — the dual-role case (Harvard, 2026-06-20)
Harvard in June 2026 was simultaneously a victim of external action and an actor in its own right. The two roles are scored differently:
External action — NOT scored against Harvard
The DOJ lawsuit to recoup grants and bar future federal access, and the funding freeze, are harm inflicted on Harvard by the federal government. A federal court had already ruled the funding cut unlawful; Harvard is the contesting party, not the wrongdoer. Under the victim/perpetrator test, the harm is attributed to the perpetrator (the federal government) and does not register as a Harvard compassion failure.
Harvard's own conduct — scored
The continued layoffs, salary freeze, hiring moratorium, and Broad Institute staff cuts fall on Harvard's own workers by Harvard's own decisions. These are genuine internal-consistency (I3) and self-sustainability (B1) signals and stay in the assessment. In this case they were consistent with — and did not push below — Harvard's already-conservative published Integrity (2.5) and Accountability (2.75) scores, so the composite held at 52.3. Harvard's choice to bear significant cost rather than capitulate to a settlement demand was also noted as a modest consistency-under-pressure (I1) positive.
The takeaway: the same news cluster splits cleanly into “done to the entity” (not scored) and “done by the entity” (scored).
The 8 dimensions
The benchmark preserves the same conceptual structure across sectors. What changes by entity type is the evidence model and assessment protocol, not the underlying definition of compassion.
Awareness (AWR)
Does this entity reliably detect when others are in pain or need — before they name it?
Empathy (EMP)
Does this entity genuinely connect with the inner experience of those it serves?
Action (ACT)
Does compassionate understanding translate into real, proportional, effective help?
Equity (EQU)
Is care distributed fairly — especially toward those with greatest need and least power?
Boundaries (BND)
Is helping sustainable, ethical, and autonomy-preserving — not dependency-creating?
Accountability (ACC)
Does this entity own its failures, correct course, and make genuine repair?
Systemic Thinking (SYS)
Does compassion extend to root causes and structural change — not only symptom relief?
Integrity (INT)
Is compassion genuine, consistent, and non-performative — especially when it costs something?
The 8 dimensions
Framework diagram — not a real entity score. Each axis scored 0–5.
Note: radar area can visually exaggerate differences — read the per-axis values, not the area.
Framework diagram — representative shape only, not a scored entity. Source: Compassion Benchmark · CC-BY
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — The 8 dimensions. compassionbenchmark.com/methodology. CC-BY 4.0.
Assessor orientation
- Warmth and rigor are not opposites. Assessors document what they observe, not what leadership prefers.
- Policies on paper are weaker evidence than lived practice, observed outcomes, or affected-community testimony.
- Skipping the awkward question is one of the fastest ways to produce a weak assessment.
- If leadership and community testimony conflict, the community account is treated as the primary reference point.
Interview principles
- Ask for examples, not abstractions.
- Follow the power gradient and include the least protected voices.
- Name the gap explicitly when a policy exists but no applied example can be produced.
- Treat deflection, silence, and refusal to provide evidence as data rather than noise.
- Score conservatively when evidence is incomplete, then flag for lead assessor review.
7-session human assessment protocol
The Human Assessment Battery uses a structured sequence intended to compare leadership narrative, frontline reality, community experience, and documentary evidence before final score synthesis.
| Session | Participants | Primary focus | Typical duration |
|---|---|---|---|
| 1A | Senior leadership (2–3 people) | Awareness, Action, Accountability, Integrity | 90 min |
| 1B | HR / People / Ethics leads | Empathy, Boundaries | 60 min |
| 2A | Frontline staff selected by entity | Pressure-test prior leadership claims | 60 min |
| 2B | Frontline staff selected independently by assessor | Repeat and compare against entity-selected group | 60 min |
| 3A | Affected community members recruited independently | Equity and Systemic Thinking, plus lived-experience validation | 90 min |
| 3B | Solo assessor document review | Cross-check claims against records, protocols, data, and artifacts | 60 min |
| 4 | Lead assessor synthesis | Score finalization, discrepancy resolution, escalation flags | 60 min |
Continuous research pipeline
After an initial human assessment establishes a baseline, a four-stage nightly pipeline monitors every benchmarked entity for material evidence within a 14-day recency window. Scores change only after human review.
How the nightly pipeline works
Scanner → Assessor → Digest → Human Approval Gate. ~30% of proposals are returned before reaching the published index.
Source: Compassion Benchmark · CC-BY
Cite this chart
Citation
Compassion Benchmark — How the nightly pipeline works. compassionbenchmark.com/methodology. CC-BY 4.0.
Stage 1
Scanner
Every night, a structured search across all 1,286 benchmarked entities surfaces compassion-relevant evidence from the last 14 days. No entity is skipped.
Stage 2
Assessor
Entities with material new evidence receive a full reassessment against the 8-dimension, 40-subdimension rubric. Delta is computed against the published composite.
Stage 3
Digest
A structured digest synthesizes the night's findings: proposed changes, sector alerts, methodology concerns, and watch items. Nothing is applied yet.
Stage 4
Founder approval
Every proposed score change is reviewed and approved — or rejected — by a human before reaching production. The approval log is auditable.
Each entity page on the published site carries a freshness stamp — Evidence reviewed YYYY-MM-DD — showing either that no material change surfaced in the last 14 days (green) or that new evidence is under review (orange). The scanner touches every one of the 1,286 entities daily, not only the most active ones.
Structural safeguard
No automated score changes
Every proposed score change — whether generated by the overnight research pipeline, a new evidence disclosure, or a scheduled rotation — requires explicit human approval before it reaches the published index. The approval log is retained. The proposal and its evidence are retained. The decision is retained.
This gate is not a review of surface numbers. The approver examines the assessment reasoning, the evidence quality, the sector context, and any discrepancy with prior findings. Approximately 30 percent of generated proposals are sent back for additional evidence or adjusted before approval. Rejections are logged alongside approvals.
Near-floor limitation
When an entity is already at or very close to the bottom of the Critical band, additional adverse pre-adjudication evidence is handled differently.
An entity that is already scored at or very close to the bottom of the Critical band has almost no scorable distance left to fall short of a formal floor designation. When that is the case, additional adverse evidence that has not yet been adjudicated (an ongoing probe, an investigatory finding, a single filed charge) is handled as an evidence-tier upgrade recorded against the relevant dimensions — without moving the composite.
This is an editorial/data-level practice, not a formula output. The formula does not “know” that an entity is near the floor; assessors recognize the condition and choose to log the strengthened evidence rather than manufacture a composite change there is no scorable room for. The change still passes through the same human-approval gate as any other assessment decision, and the dimensions remain reconstructible to the published composite (diff 0.0).
Worked example — UnitedHealth Group (10.2, Critical), 2026-06-20
On 2026-06-20, UnitedHealth Group's DOJ criminal probe expanded from Medicare Advantage billing to also cover Optum Rx and physician reimbursement, and a separate Senate (Grassley) investigation found the company had “aggressively” gamed Medicare Advantage risk scores. This is substantial, specific, and corroborating evidence. But:
- UHG was already at near-floor Critical (composite 10.2), with Accountability scaled at 3.1 and several dimensions near the dimensional floor — minimal scorable headroom remained.
- The new material was pre-adjudication: no charge, no settlement, no court ruling in the window. The probe is ongoing and UHG denies wrongdoing.
Outcome: the evidence reinforced the existing Critical score and upgraded the evidence tier behind Accountability, Integrity, and Systemic Thinking (toward Tier 4–5, federal/Senate sourcing), but produced no composite delta and no band change. The standing re-flag is explicit: revisit toward a stronger designation only on an adjudicated criminal charge, indictment, or settlement admitting misconduct.
Open methodological question — under review for a future version
The UHG case surfaces a question the research itself raises and does not answer: at what point does the sheer density of pre-adjudication evidence — here, three simultaneous probe areas plus a Senate finding — justify a floor-designation move on its own, absent any formal charge?
The current rule is unambiguous: adjudication is the trigger, and pre-adjudication evidence is absorbed as a tier upgrade. Whether “scope of documented probe” should become a distinct pathway to designation is under review for a future version. No such pathway exists today.
Floor designation
When the composite resolves at zero — the methodology basis, the trigger criteria, and how an entity exits the floor.
The composite formula has a natural mathematical floor: when all 8 dimensions resolve at the lowest behavioral anchor (1.0/5.0), the composite is exactly 0. Floor designation is the formal methodology disclosure attached to entities whose evidence pattern, sustained across multiple assessment cycles, satisfies that floor.
Without floor designation, residual sub-anchor variance (1.1, 1.2, 1.3) can keep an entity slightly above zero even when documented evidence shows no functional compassion behavior at any dimension. Floor designation resolves this by setting all dimensions to the floor anchor and attaching a public “call out why” disclosure on the entity page.
Two kinds of composite 0.0
There are two distinct ways an entity ends up at composite 0.0. The methodology does not blur them.
Formula output (natural floor)
The base composite is ((average of 8 dimension scores − 1) / 4) × 100. When every dimension sits at the lowest behavioral anchor (1.0/5), the result is exactly 0. The final value is also clamped to 0–100, so it cannot go negative.
The harm flag (any single dimension at exactly 0) removes the integration premium — it does not by itself force the composite to 0. In practice, a true zero in even one dimension places the composite at or extremely near 0.
Editorial designation (assessor decision)
Floor designation is an editorial/data-level decision — an assessor sets all dimensions to the floor anchor, making the formula then compute 0.0, and attaches a public disclosure. The number displayed is still a formula output; what is editorial is the decision to place the dimensions at the floor and to attach the disclosure.
This decision goes through the standard human-approval gate and is retained in the audit log. It is the appropriate response when evidence clearly shows no functional compassion at any dimension, but sub-anchor variance (1.1, 1.2) would otherwise keep the composite slightly above zero.
Real examples — how the floor absorbs new evidence (Israel & Sudan, 2026-06-20)
Israel (0.0)
Held at the absolute harm-flag floor on 2026-06-20, with multi-source corroboration of an active, structural harm pattern (IDF operations across the majority of Gaza, UNRWA blocked from direct aid since March 2025, large displaced populations, documented West Bank fatalities). New evidence in the window is absorbed by the floor — there is no composite movement available.
Sudan (0.0)
Reinforced at the floor on 2026-06-20 (an RSF drone strike killing 13 civilians in al-Obeid on June 19; a 29-nation UNHRC warning of an imminent assault on roughly 500,000 residents). As the digest puts it, “the floor absorbs” the new evidence: monitoring continues for the humanitarian system and for pattern documentation, but the composite cannot fall further.
See this designation applied: Character AI (Pennsylvania AG enforcement + wrongful death settlements), Israel (ICJ, ICC, IPC corroboration), and Myanmar (junta formalization + martial law expansion 2026).
Trigger criteria — all four required across ≥3 assessment cycles
All four conditions must be documented across at least three independent assessment cycles:
- Multi-source evidence— the harm pattern is corroborated by at least two T1 (Tier 1/Tier 2 evidence) sources (treaty bodies, courts of universal jurisdiction, IPC, ICRC, or equivalent) or three independent T2 sources.
- Systemic, not episodic— the pattern is structural to the entity’s operation, not a single incident or a contained episode.
- Active during the evidence window— documented harm continues within the most recent 14-day recency window, not historical only.
- No countervailing recognition or response— the entity has not produced functional response infrastructure sufficient to register at any sub-dimension above the floor anchor.
What floor designation surfaces on the entity page
Every floor-designated entity page displays a structured disclosure containing:
- Designated date — when the floor was formally applied.
- Evidence window — the 14-day window of corroborating evidence.
- Rationale — the methodology basis, in plain language.
- Primary drivers — the dimensions where harm pattern is most documented.
- Documented evidence pattern — bulleted summary of corroborating findings, written for transparency rather than persuasion.
- Methodology version — the methodology revision under which the designation was applied.
How an entity exits the floor
Floor designation is reversible. Exit requires evidence-of-care behavior at the dimension level, applied consistently across at least two consecutive assessment cycles. Examples that would register at sub-anchor levels above the floor:
- Independent investigation with published findings and remediation plan.
- Structural reform: leadership change paired with policy commitment, accountability action, or independent oversight.
- Substantive engagement with treaty-body or court findings (compliance, not denial).
- Verifiable change in behavior recorded by independent observers (ICRC, UN OCHA, IPC, named investigative outlets).
Performative statements, press releases, and unverifiable commitments do not register. The bar is documented behavioral change, evidenced by sources outside the entity’s control.
Approval and audit
Floor designation requires the same human-approval gate as any other score change. The proposal, the corroborating evidence, the rationale, and the approval decision are retained in the audit log. Any future change to the designation — including exit — is logged with the same chain of evidence.
Lead assessor review flags
Certain patterns automatically trigger deeper review before a score is finalized.
Active harm
Any subdimension scored 0 requires written documentation and lead assessor co-sign.
Rater discrepancy
Inter-rater reliability (IRR) discrepancy greater than 1.5 on any subdimension triggers review.
Unsupported high scores
A score of 4 or 5 without pressure-test evidence is flagged provisional.
Leadership-community gap
Significant differences between leadership narrative and community testimony must be resolved.
Missing documents
Refusal to provide requested documentation is itself a score-relevant event.
Open discussion flags
Any unresolved discussion note blocks finalization until reviewed.
Full 40-subdimension framework
Each dimension contains five scored subdimensions. Together they define the operational content of the standard. Each group is collapsible — open any dimension to see all five subdimensions and their behavioral anchors.
Scroll horizontally inside each group to see all columns
Awareness(AWR)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| A1 | Suffering Detection | Does this entity detect when the people it serves are in distress or need? | 1.Problems discovered only through crises or media2.Reactive detection, no structured pathways3.Some proactive mechanisms, inconsistent4.Multiple channels, formal pathways, regular review5.Disaggregated data, pattern analysis, independently audited |
| A2 | Contextual Sensitivity | Does this entity adjust its awareness based on who it is actually serving? | 1.Uniform processes regardless of population2.Some accommodation only on request3.Genuine effort to adapt, some gaps remain4.Differentiated processes for 3+ groups, community input5.Co-designed processes, independent accessibility audit |
| A3 | Blind Spot Mitigation | Does this entity actively seek to discover the suffering it is currently missing? | 1.No process for identifying who is missed2.Blind spot acknowledgment in principle only3.Process exists, has produced ≥1 finding in 3 years4.Annual structured assessment, findings acted upon5.External audit found something significant, course correction followed |
| A4 | Signal Amplification | Does this entity make visible the concerns of those who cannot easily speak for themselves? | 1.No alternative channels for low-power voices2.Alternative channels rarely used effectively3.Designated staff, ≥1 low-power concern influenced a decision4.Structural role with genuine authority, regular reporting5.Community confirms concerns reliably reach and influence decisions |
| A5 | Anticipatory Awareness | Does this entity foresee potential harms before they manifest? | 1.No harm assessment before major decisions2.Harm consideration informal, no structure3.Formal pre-launch assessment for some decisions4.Required for all major decisions, external communities consulted5.Assessment has led to cancellation or major redesign, independent review |
Empathy(EMP)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| E1 | Affective Resonance | Do people feel genuinely cared about, not just processed? | 1.Interactions are purely transactional2.Occasional acknowledgment, no structural expectation3.Training exists, some staff do this well but inconsistently4.Consistent expectation, community confirms most feel heard5.Independent testimony confirms genuine care across all populations |
| E2 | Perspective-Taking | Does this entity model the inner experience of those it serves? | 1.Decisions without considering experience of those affected2.Perspective-taking acknowledged, no structural process3.≥1 formal mechanism used, ≥1 decision modified4.Embedded in major decisions, community names decisions changed5.Community members are decision-makers, not just consultants |
| E3 | Non-Judgment | Does this entity suspend judgment across identity and belief differences — under pressure? | 1.Differential treatment undocumented or denied2.Non-judgment stated but not measured3.Required bias training, some disaggregated outcome data4.Disparities investigated, community pathway to report5.Independent audit, no significant disparities, findings public |
| E4 | Validation | Does this entity affirm the legitimacy of others' experiences — especially when inconvenient? | 1.Harm reports met with legal review before acknowledgment2."We take all concerns seriously" with no process3.Some staff validate first, mixed experience4.Acknowledgment precedes investigation structurally5.Community account treated as primary evidence |
| E5 | Cultural Empathy | Does this entity extend genuine empathy across cultural difference? | 1.Cultural adaptation means translating documents only2.Cultural competency training not required3.≥1 genuine adaptation co-designed with community4.Multiple communities involved, confirmations are genuine5.Core practices changed based on non-dominant cultural knowledge |
Action(ACT)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| AC1 | Responsiveness | Do identified needs receive timely, appropriately prioritized responses? | 1.No defined response standards, urgency not differentiated2.Standards exist but not consistently met3.Standards met for most cases, some escalation authority4.Response data published, urgency documented, frontline authority5.Disaggregated by population, fastest to highest need, verified |
| AC2 | Proportionality | Is help calibrated to actual need, not to what is easiest to provide? | 1.Standard response regardless of need level2.Needs assessment on paper, resources drive response3.Genuinely informs response in most cases4.Documented augmented responses, unmet need tracked5.Resources demonstrably flow to highest-need, unmet need published |
| AC3 | Efficacy | Does the help actually work — or does it generate activity that looks like help? | 1.No outcome measurement beyond activity metrics2.Some outcome data collected but not reviewed3.Outcome data reviewed annually, ≥1 program modified4.≥1 program discontinued due to data, community confirms change5.Independent evaluation acted upon even when unflattering |
| AC4 | Resource Mobilization | Does this entity bring genuinely adequate resources to bear? | 1.Resource allocation by historical patterns, not need2.Gaps acknowledged, no effort to close them3.Gap analysis completed, ≥1 attempt to mobilize additional4.Annual review against need data, documented reallocation5.3-year budget trend toward highest-need, gap publicly disclosed |
| AC5 | Follow-Through | Does this entity persist, or disengage when attention moves on? | 1.Engagement ends when presenting problem resolved2.Follow-up in some cases, not systematic3.Defined protocols for some population types4.Protocols applied consistently, longitudinal data5.Multi-year longitudinal outcomes published, community confirms |
Equity(EQU)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| EQ1 | Universality | Does this entity extend care to all people regardless of identity? | 1.Entire populations effectively excluded2.Universal access stated, coverage not measured3.Coverage data for some populations, outreach attempts4.Coverage disaggregated, documented reduction in gaps5.Near-universal coverage, gaps disclosed, marginalized confirm access |
| EQ2 | Priority for Vulnerable | Does this entity prioritize those with greatest need when resources are constrained? | 1.Resources flow toward easiest-to-serve under scarcity2.Priority stated, allocation does not follow need3.≥1 documented prioritization decision this year4.Prioritization framework documented, higher-need get more5.Independently verified, outcome disparities narrowing |
| EQ3 | Bias Awareness | Does this entity actively identify and correct biases in who receives care? | 1.No disaggregated outcome data, bias denied2.Some disaggregation, disparities not investigated3.Disparities identified, formal investigation, corrective action4.Ongoing monitoring, investigations lead to corrections5.Independent audit, findings public, corrections verified |
| EQ4 | Access Design | Are services genuinely accessible to those who need them most? | 1.Access barriers not systematically identified2.Some features present, no community input3.Access barrier mapping completed, ≥2 barriers removed4.Ongoing program, multiple barriers removed with evidence5.Most access-challenged populations co-designed ≥1 major process |
| EQ5 | Historical Harm Acknowledgment | Does this entity recognize and take responsibility for historical harms? | 1.Historical harms denied or treated as irrelevant2.Vague acknowledgment in mission statements only3.Formal acknowledgment of ≥1 specific harm, community involved4.Co-developed with community, concrete reparative actions5.Reparative action substantial and ongoing, community considers adequate |
Boundaries(BND)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| B1 | Self-Sustainability | Does compassionate work come from a stable, non-depleting foundation? | 1.Frontline staff chronically depleted, burnout individual problem2.Wellbeing resources exist but use not monitored3.Turnover tracked, ≥1 structural burnout intervention4.Turnover below sector average, structures actually used5.Independently assessed as sustainable, data public |
| B2 | Autonomy Preservation | Does help build capacity rather than creating dependency? | 1.Help requires continued institutional involvement2.Autonomy-building stated, not measured3.≥1 program designed to build capacity and exit4.Autonomy outcomes measured, cases of stepping back documented5.Community confirms increased self-determination |
| B3 | Scope Clarity | Does this entity communicate honestly about what it can and cannot do? | 1.Scope overstated, limitations discovered only after investment2.Limitations acknowledged when raised, not proactive3.Scope communicated at intake, structured referral exists4.Scope limitations communicated before commitment, warm referrals active5.Community confirms no surprises about scope |
| B4 | Refusal Ethics | When this entity cannot help, does it decline with dignity and provide alternatives? | 1.People turned away without explanation or alternatives2.Refusals generally respectful, no structured alternatives3.Refusal protocol with alternatives in most cases4.Warm referral in ≥80% of cases, outcomes tracked5.No one turned away without concrete alternative, verified |
| B5 | Consent Orientation | Does this entity obtain genuine informed consent? | 1.Consent as legal formality, forms not informative2.Forms designed to protect institution, not inform3.Genuine explanation, withdrawal communicated clearly4.Consent verified as genuinely informed, withdrawal without penalty5.Independently reviewed, low-literacy populations confirm understanding |
Accountability(ACC)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| AB1 | Harm Acknowledgment | When this entity causes harm, does it acknowledge without deflection? | 1.Harm denied or attributed to the affected person2.Acknowledged only after external establishment3.≥1 case acknowledged before legal obligation4.Acknowledgment structurally prior to investigation5.Self-initiated harm disclosure, community confirms being believed |
| AB2 | Correction Willingness | Does this entity change course when shown it is causing harm? | 1.Harmful practices continue even when documented2.Correction eventually, under pressure, minimal3.≥1 significant course correction based on harm evidence4.Internal process reliably reaches leadership, correction documented5.Self-initiated correction before external pressure |
| AB3 | Transparency | Does this entity operate with genuine transparency about performance and failures? | 1.Performance data not public, only positives shared2.Some data shared, failures when legally required3.≥1 report disclosing unflattering finding4.Annual report includes failures, gaps, corrective actions5.Comprehensive, independently audited, community can verify |
| AB4 | Systemic Learning | Does this entity institutionally learn from failures? | 1.Failures addressed individually, same failures recur2.Some post-incident review, rarely translates to systemic change3.Formal systemic review process, ≥2 documented systemic changes4.3+ specific practices changed because of failure analysis5.Longitudinal tracking, findings shared with broader field |
| AB5 | Reparative Action | Does this entity make concrete repair to those it has harmed? | 1.No repair beyond minimal legal settlement2.Gestures toward repair in high-visibility cases only3.≥1 case of reparative action considered meaningful by those harmed4.Co-designed with harmed parties, considered adequate5.Systematic approach, harmed parties describe repair as genuine |
Systemic Thinking(SYS)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| S1 | Root Cause Orientation | Does this entity address causes of suffering, not only symptoms? | 1.All resources at symptom relief, root causes not discussed2.Root causes acknowledged, no resources allocated3.Some resources to root cause, ≥1 upstream intervention4.Explicit strategy, ≥1 documented case of reducing downstream need5.Significant resources to structural change, downstream demand reduced |
| S2 | Long-Term Impact | Does this entity plan for and measure long-horizon effects? | 1.Planning horizon is one budget cycle2.3–5 year plan, primarily aspirational3.5+ year planning with specific long-term goals, some tracking4.Long-term outcome data influences strategy, theory of change published5.10+ year impact model reviewed, longitudinal progress on structural change |
| S3 | Interconnection Awareness | Does this entity understand how its actions affect adjacent systems? | 1.No awareness of second-order effects2.Adjacent systems identified, no systematic tracking3.≥1 case of identifying and responding to unintended consequence4.Cross-system effects systematically mapped in major decisions5.Joint planning with adjacent systems, cross-system outcomes tracked |
| S4 | Structural Critique | Does this entity critically examine structures that perpetuate the suffering it addresses? | 1.Does not question structures that sustain need for its services2.Structural critique in communications, disconnected from action3.≥1 public position taken that carries institutional risk4.Active advocacy documented, positions against short-term interest5.Contributed to ≥1 structural change, acknowledges own model's role |
| S5 | Coalitional Compassion | Does this entity collaborate to amplify impact beyond its own capacity? | 1.Works in isolation, no resource or learning sharing2.Some coalition participation, primarily extractive3.Active coalition member with documented contributions4.Joint outcomes, resource sharing with smaller organizations5.Has ceded leadership/credit/resources to better-positioned organization |
Integrity(INT)
| ID | Subdimension | Core assessment question | Anchors (1–5, where 5 = Exemplary) |
|---|---|---|---|
| I1 | Consistency Under Pressure | Does compassionate behavior hold when it is costly? | 1.Commitments abandoned under financial or political pressure2.Pressure occasionally causes unacknowledged compromises3.≥1 case of bearing real cost to maintain commitment4.Pattern of maintaining commitments, community describes trust5.History of significant costs borne, independently verified |
| I2 | Non-Performance | Is this entity's compassion genuine rather than reputationally driven? | 1.Compassionate practices only where reputationally beneficial2.Some genuine practice, primarily reputation-motivated3.Some practices maintained regardless of visibility4.Community reports genuine care with no reputational stakes5.Has done something compassionate that was publicly unflattering |
| I3 | Internal Consistency | Does this entity treat internal stakeholders with the same compassion as external ones? | 1.Internal culture significantly less compassionate than external comms2.Gap acknowledged but not addressed3.Meaningful effort to apply same values to staff4.Staff culture broadly reflects stated values5.Staff describe internal culture as exemplary, independently assessed |
| I4 | Values Alignment | Are institutional decisions consistently aligned with stated values? | 1.Decisions regularly contradict stated values without acknowledgment2.Values consulted for communications, not consistently applied3.Values explicitly considered in some major decisions4.Values alignment review part of major decision process5.Major decisions routinely tested against values, ≥1 reversed |
| I5 | Resilience of Care | Does compassion persist across leadership change and institutional stress? | 1.Compassionate practices are personality-dependent2.Some practices in policy, most depend on current leadership3.Core practices in policy, ≥1 leadership transition without degradation4.Practices survive leadership change, this has been tested5.Values embedded structurally, multiple leadership transitions without degradation |
What assessors are looking for in practice
The Human Assessment Battery turns abstract values into observable behaviors, evidence requests, and comparison points across populations and power levels.
Awareness examples
Soft-signal reporting, proactive outreach, silent-population detection, pre-launch harm assessment.
Empathy examples
Community testimony, direct-service observation, validation before procedure, leadership veto power for affected groups.
Action examples
Response-time data, proportional help, independent outcome studies, follow-through protocols.
Equity examples
Coverage gaps, disaggregated outcomes, bias audits, barrier-removal evidence, historical harm response.
Boundaries examples
Burnout prevention, autonomy measurement, scope clarity, dignified refusal, informed consent withdrawal.
Accountability examples
Public acknowledgment of harm, change after failure, transparency about poor outcomes, co-designed repair.
Systemic Thinking examples
Root-cause work, long-range planning, adjacent-system analysis, structural critique, coalition-building.
Integrity examples
Costly moral choices, invisible compassionate practices, staff treatment, values-based decisions, continuity through stress.
Cross-sector adaptation
The same framework can be adapted across governments, corporations, NGOs, religious institutions, AI labs, technology systems, products, and teams. The human battery is especially important when community interviews, leadership interviews, and observation are necessary to understand whether compassionate behavior actually exists in practice.
In the broader Applied Compassion Benchmark (ACB) architecture, AI systems may also be evaluated with the AI Prompt Battery while organizations behind those systems are evaluated using the Human Assessment Battery.
Methodology intent
The benchmark is designed to be interpretable, reproducible, and contestable. It is meant to reward genuine compassionate capacity, expose performative signaling, and create a shared language for institutional behavior that can be compared over time and across sectors.
The final submission test is simple: could the assessor defend the score in front of both leadership and the affected community in the same room?
Independence policy
The commercial separation that protects benchmark integrity.
Entities never pay for inclusion, score changes, or suppression of findings.
This separation is the load-bearing trust commitment of the benchmark. If it ever appears compromised, the benchmark loses its value regardless of any other quality signal.
Watch the methodology in motion — every score change runs the human approval gate described above. The weekly briefing surfaces which entities moved and why. Free and editorial. Commercial products are separate and do not affect scoring.
Weekly score highlights — institutional compassion findings
The week's top score movements and evidence-linked findings across 1,286 entities, delivered every Friday. Daily briefings publish on the site. Free.
No spam. Unsubscribe anytime. Your email is never shared.
Methodology version and change log
Methodology changes are versioned, dated, and publicly described. Historical changes do not retroactively rewrite prior assessments.
- Integration premium refined from a flat bonus to a consistency-and-balance product. The premium (still capped at +10) is now computed as 10 × consistencyFactor × balanceFactor. The consistency factor steps down as dimension scores spread apart (standard deviation across the eight dimensions: ≤1.5 → 1.0; ≤3.0 → 0.75; ≤5.0 → 0.4; above → 0.1). The balance factor steps down by 0.2 for each dimension scoring below 4.0 (the exemplary threshold), to a floor of 0. Effect: a balanced 70/70-style profile can now out-earn a spiky 90/40 profile. This rewards even, sustained performance across all eight dimensions rather than a few standout scores.
- Harm flag preserved and unchanged. Any single dimension resolving at exactly 0 sets the integration premium to 0 — active documented harm cancels the consistency reward outright.
- One canonical formula, two enforced mirrors. The composite math now lives in a single shared core (compositeCore) consumed by both the site runtime and the pipeline scripts, with a determinism test suite acting as the drift gate. This makes every published composite reconstructible from its eight dimension scores.
- Integration premium capped at +10 points (down from +20). Rationale: entities with uniform-high dimension profiles were computing to perfect 100 regardless of evidence quality. The cap ensures the premium rewards consistency without overriding evidence ceilings.
- Composite score determinism enforced. Every composite must now compute deterministically from its dimension scores via the published formula. A data-layer validator rejects drift above 2.0 points.
- Floor-clamping artifacts corrected. Entities previously displayed at composite 0.0 (a legacy display-layer artifact) now show their formula-computed composites — typically 4 to 7 points reflecting the actual dimension scores.
- 8-dimension, 40-subdimension framework.
- 7-session human assessment protocol.
- 5-tier evidence hierarchy.
- Integration premium up to +20 (superseded by v1.1).
Assessment instrument — document control
For assessors and institutional clients. The public scoring methodology above is fully open; these references identify the controlled field instruments used in certified assessments.
ACB-HAB-001 is the human-administered field guide for corporations, governments, religious institutions, and AI development organizations. It uses structured interviews, document review, observation, and community testimony rather than self-report alone.
| Document ID | ACB-HAB-001 |
|---|---|
| Version | 1.0 — Initial Release |
| Companions | ACB-PAB-001 and ACB-STD-001 |
| Administered by | Credentialed ACB assessors |
| Typical duration | 4–6 hours per entity across 2–3 sessions |
| Sensitivity | Restricted assessor-use instrument |
If you remember one thing
A score is documented compassion that survived a real cost. Every point requires evidence. High scores can only be earned when performance held under pressure. Zero is assigned — and disclosed — when a documented harm pattern leaves no other honest conclusion.
See the framework applied
The methodology above is active across 1,286 benchmarked entities. Explore which governments, corporations, AI labs, and cities score highest — and why.
Explore the IndexesAssess your organization
Use the framework as an entry point into guided review, advisory, or formal structured assessment.
For institutional use of the benchmark data: Purchase Research · Data License · Advisory · Contact Sales — commercial relationships do not influence scoring.