The high-level float had been failing intermittently since installation. It took three weeks of a permanent fault for anybody to look, and the log looked normal throughout.
Why the log looked fine
The alarm was configured to raise on a state change and clear on the opposite one. An intermittent contact produced pairs of raise-and-clear events seconds apart, which the summary view counted as resolved alarms. By any measure the operators had, the station was healthy.
Nothing was recording the thing that was actually drifting: the number of transitions per day. That figure had been climbing for eight months.
What an alarm is for
An alarm that clears itself is not information, it is noise with a timestamp. The useful signal was never the state — it was the rate of change of the state, and no threshold had been set on it because nobody had thought to.
What changed
A derived alarm on transition count, and a rule that any alarm clearing in under sixty seconds is logged as a chatter event rather than a resolution. The first week it ran, it found two more floats in the same condition at other stations.
