Most AI benchmarks like SWE-bench or Terminal Bench mean nothing to the average person. BrushArena makes model progress legible: models write code that renders into paintings, so anyone can judge a model’s capabilities for themselves.