Research question How should severity language be calibrated so empty findings never become false assurance?