AI Thinks Like Us: ChatGPT Shows Overconfidence and Human Biases

Although humans and artificial intelligence systems think very differently, new research shows that AI often makes irrational decisions—just like we do.

Researchers from five academic institutes in Canada and Australia tested two large language models (LLMs) – GPT-3.5 and GPT-4 from OpenAI. The team found that despite their “striking consistency” in reasoning, the models are not immune to human flaws.

In almost half of the scenarios examined by the new study, ChatGPT demonstrated many of the most common human biases in decision-making.

This work is the first to evaluate ChatGPT’s behavior across 18 well-known cognitive biases identified in human psychology. The study looked specifically at widely recognized biases such as risk aversion, overconfidence, and the endowment effect (when we assign extra value to our possessions). The researchers built these bias scenarios into ChatGPT prompts to see whether it would fall into the same traps as humans.

What did the experts find out?

Scientists posed hypothetical questions to the language models that were borrowed from traditional psychology, but placed in real commercial contexts such as inventory management and negotiations with suppliers. Researchers didn’t just want to see whether AI would copy human biases; they also wanted to know whether it would do so when answering questions from different business sectors.

GPT-4 outperformed GPT-3.5 on tasks with clear mathematical solutions, making fewer errors on probabilistic and logical problems. However, in subjective simulations—such as choosing a risky option for profit—the chatbot often behaved irrationally, like people do.

Like people, AI systems are overconfident and biased.

Researchers also noted that the models tend to prefer safer, more predictable outcomes when faced with ambiguous tasks.

The models’ behavior remained largely stable regardless of whether the questions were framed as abstract psychological problems or as operational business decisions. The researchers concluded that the biases they found were not merely examples the models had memorized, but were tied to the models’ underlying reasoning patterns.

One surprising result was that GPT-4 sometimes amplified human mistakes. The authors wrote in the report that, in the biased-confirmation task, GPT-4 consistently provided biased responses. Live Science highlighted that finding and also reported that GPT-4 showed a stronger tendency toward the hot-hand fallacy—the bias of expecting patterns in randomness—than GPT-3.5.

On the other hand, ChatGPT avoided some common human biases, particularly ignoring fundamental metrics (favoring anecdotes over statistics) and the sunk-cost fallacy (letting past costs improperly influence current decisions).

Researchers say these “human” biases in ChatGPT arise because the training data contains the cognitive biases and heuristics typical of people. Those tendencies can be amplified during fine-tuning, especially when human feedback rewards responses that seem plausible rather than strictly rational. When models face ambiguous tasks, they are more likely to lean on human-like reasoning instead of direct logic.

“If you need accurate, unbiased decision-making support, use GPT in areas where you already trust a calculator,” advised lead researcher Yan Chen. He added that when outcomes depend on subjective or strategic inputs, human oversight is essential, including adjusting prompts to correct known biases.

“AI should be treated as a colleague making important decisions; it requires oversight and ethical principles. Otherwise, we risk automating flawed thinking instead of improving it,” said Mina Andiapann, a co-author of the study.

The results of the study were published in the journal Manufacturing & Service Operations Management.