A paired A/B benchmark of the token-compression skill Caveman on Claude Code, run on SkillsBench: does it actually save tokens, and does it degrade AI agent output quality?
Advertised saving: 65%. Measured saving: 8.5%.
Output-token saving on real agentic tasks, with the skill forcibly activat…