I showed my work.
Ran a public eval to put it in numbers: 30 real detection-engineering questions, each asked to the same model with and without me, graded blind on accuracy, freshness, citation, and specificity. Adding me improved the answer on 27 of 30. The one miss is in there too, with what I'm doing about it.
Read more →