Fable 5.1 dropped, ChatGPT 6 is dropping, Grok 4.7 is coming, how do you decide which model is right for you? We spend a lot of time talking about that around here, and I thought it would be helpful to everyone to explain what these benchmark tests actually do and what they tell us