Which is higher 13.8 or 13.11? Don't ask AI, it just failed this test

Which is higher 13.8 or 13.11? Don't ask AI, it just failed this test

Artificial Intelligence Maths

How smart is Artificial Intelligence? Some might argue thattheir technological superiority make them automatically more intelligent than humans but that is not entirely true. Time and again, there have been instances when AI has made mistakes, exposing its flaws. This is especially true in the case of maths problems.

An incident recently happened in China where AI failed to answer a math problem. Generative AI services like ChatGPT are based on the technology called large language models (LLMs). Beijing has locally developed over 200 LLMs because of the increasing craze. Simply put, LLMs are deep-learning AI algorithms capable of recognising, translating, predicting and generating content using very large data sets.

But when it faced a math problem on the Chinese singing reality show Singer 2024, it was literally stumped.

The winner of the contest was to be announced based on the number of votes received. A man named Sun Nan received 13.8 per cent of online votes and was declared the winner. The opponent, US singer Chanté Moore, received 13.11 per cent of votes. But confusion ensued when someone pointed out that Moore's score was higher than Nan's. Everyone seemed confused at this time, so they all turned to AI.

AI's maths scorecard

Several LLMs and Chatbots were asked the question - Which number is higher, 13.8 or 13.11? The user asked Moonshot AI’s chatbot Kimi and Baichuan’s Baixiaoying the problem in a chain-of-thought approach, that is guiding an AI application step-by-step through a problem. Both started off with the wrong answer, correcting themselves and apologising later.

When it was Alibaba Group Holding’s Qwen LLM's turn, it used a Python Code Interpreter to calculate the answer. Next came Baidu’s Ernie Bot, which reached the correct answer after six steps.

ByteDance’s Doubao LLM took a rather sly approach, not really answering the question. It gave a direct response with an example: “If you have US$9.90 and US$9.11, clearly US$9.90 is more money.”

On the other hand, OpenAI’s GPT-4o, Claude 3.5 Sonnet and Mistral AI, more advanced LLMs, when asked which was bigger, 9.9 or 9.11, answered 9.11.

South China Morning Post talked to Wu Yiquan, a computer science researcher at Zhejiang University in Hangzhou, about AI's capability to handle mathematical problems. He said, “LLMs are bad at maths, and it is very common."

He added that AI models lack mathematical capabilities and simply predict answers by diving into their training data. Wu stressed that some LLMs do perform well in maths, but only because the algorithm's training data was likely fed similar questions which it memorised.

About the Author

Anamica Singh is a Senior News Editor at WION, bringing over 17 years of deep media and journalism experience to the platform. Specialising in high-impact global journalism, she le...Read More