Apple draws attention to a persistent problem in language models: their reliance on pattern matching rather than genuine logical reasoning. In several tests, the researchers demonstrated that adding irrelevant information to a question—details that should not affect the mathematical outcome—can lead to vastly different answers from the models.

Source: Michael Tsai – Blog – Understanding the Limitations of Mathematical Reasoning in Large Language Models