Centered Care
Every warehouse object the Superset dashboards read, checked for gone-quiet communities, missing months, broken trends, impossible values, duplicate keys and column shape. Read-only. One run per row.
sql/20260904_data_quality_schema.sql from the
cc-superset-table-validation skill in the Supabase SQL editor.data_quality. This step is project-global and cannot be done
from SQL.python scripts/publish.py <validation_*.json> --send.The weighted roll-up of the three quality dimensions below. Coverage is the fourth dial and is deliberately not folded in: one number cannot say both "the data is good" and "we could check most of it".
One point per published run, per dimension.
Quality first, then how much of the estate that check reached. A check can score well simply because it barely ran, which is what the second number is for.
Ordered by weighted points lost, not by how many findings there are. One P0 on a table a PROD dashboard reads outranks forty P2s on something nothing reads.
Two collapses. A community behind in several objects at once is one feed or DAG problem. The same finding at many communities of one object is one table defect, not one per community. Both are filed once here, and neither changes the score.
Nothing here knows which communities are supposed to feed which object. A community that never enrolled in a care program correctly has no row in that program's table. So when one pattern covers most of an object's communities, that is evidence about the object's scope rather than about each community, and it is filed once, here, instead of one finding per community. Each row still needs a person to say "scoped" or "broken". These do not affect the score.
Everything not folded into a pattern above. Filter, then read the sentence: each one says what is wrong and why it matters.
Weighted by severity, then by how many objects the community appears in. Near the top of this list across many objects means one upstream cause, not many table problems.
Off-boarded communities, each with the decision that justifies it. Never
inferred from silence: gating on dim_tenants.active alone
once closed eight live communities and 757 residents.
The most important table on the page. A report that lists only findings implies everything else passed. Each row says what could not be checked and what that costs.
How each object was read. Grain type decides which checks even make sense, and getting it wrong is the biggest source of false positives.
The last automated Centered Care data-quality signal, a daily audit email, was switched off in August 2026 for being noise. A single unweighted score is that same failure with a number attached. So quality only counts checks that ran, and coverage carries the rest.
Each (object, check) pair is one cell. A cell's credit starts at 1.0 and
loses (1 - severity credit) x (0.4 + 0.6 x share of communities
affected). Severity sets how big the penalty is, breadth sets how
much of it lands, and the 0.4 floor is there so a P0 at a single
community is never free. The cell is then weighted by how much the check
matters and by how many things read the object.