
Testing Models Against a Real Use Case, Not Just Benchmarks
·9 mins
Blind-testing two Ollama models against real CIPLE A2 Portuguese-exam questions, the well-documented, benchmark-topping model lost badly to an obscure community fine-tune — because the benchmarks were measuring Brazilian Portuguese, not European.