Which architectural claims survive at small scale under a fair, tuned baseline. One number, matched compute, negative results kept.
Seven efficient-architecture papers beat a matched transformer at tiny scale. Give the baseline its own best learning rate and double the width, and six of seven wins collapse into a tie — with a clean gated-recurrence lean.
Read →