Auditing a rule base¶
A fuzzy system is supposed to be readable. That does not stop a rule base — hand -written or machine-learned — from contradicting itself, repeating rules, leaving regions where nothing fires, or carrying terms so similar that no human can tell them apart.
None of this is visible at the call site. The system still returns a number.
audit is what makes it visible.
Running an audit¶
from fuzzytool import datasets
from fuzzytool.audit import audit
system, *_ = datasets.credit_risk()
print(audit(system))
audit: 3 rules, coverage 91.7%
unused terms (1):
dti[low]
uncovered inputs (10):
{'score': 630.0, 'dti': 0.0}
{'score': 630.0, 'dti': 5.0}
...
That is this library's own demo system, and the audit finds a genuine hole in
it: a borrower with a mid-range score (630-685) and low leverage matches no
rule, so the output falls back to the on_no_rule policy rather than to
anything the author intended. The term dti[low] is defined and never used.
Neither fact is apparent from reading the three rules.
What it looks for¶
| Finding | Meaning |
|---|---|
contradictions |
two rules share an antecedent but disagree on the consequent |
duplicates |
a rule repeats an earlier one verbatim |
never_fired |
a rule that fires nowhere in the sampled input space |
unused_terms |
a term no rule mentions |
indistinguishable |
two terms of one variable overlapping past a threshold |
partition_gaps |
points of an input universe where every term is zero |
uncovered / coverage |
inputs where no rule fires at all |
Coverage is measured by sweeping a lattice over the input universes
(n_per_axis, capped at max_points), so it is an estimate — a narrow but
valid rule can look dead under a coarse sweep. That is exactly why prune does
not remove dead rules unless you ask.
Everything works on every engine: Mamdani, TSK, Tsukamoto, FuzzyClassifier
and the interval type-2 engines all expose the same rules/firing interface.
Pruning¶
from fuzzytool.audit import prune
lean = prune(system) # drop verbatim duplicates
lean = prune(system, drop_never_fired=True) # also drop dead rules
lean = prune(system, min_weight=0.1) # and rules nobody trusts
prune returns a new system of the same class with the same configuration,
holding the surviving rules in their original order. It never mutates the
original, so you can compare the two before committing to the shorter one.
This matters most after rule learning: Wang-Mendel and Chi both emit one rule per occupied partition cell, and a good fraction of those are duplicates or near-dead once you look.
Interpretability metrics¶
{'n_rules': 3, 'n_conditions': 6, 'mean_rule_length': 2.0,
'max_rule_length': 2, 'n_input_variables': 2, 'n_terms': 7,
'coverage': 0.917}
Report these next to accuracy. A rule base that is 2% more accurate with four times the rules is usually the worse model, and this is the number that says so.
Comparing terms directly¶
The audit builds on fuzzytool.measures, which you can
also use on its own:
from fuzzytool import measures as ms
x = score.universe
ms.jaccard(ms.sample(score.terms["fair"], x), ms.sample(score.terms["good"], x))
ms.consistency(...) # can these two terms ever hold at once?
ms.subsethood(...) # is one term contained in the other?
See also: Rule learning, Measures & relations.