Anthropic has developed an AI 'brain scanner' to understand how LLMs work and it turns out the reason why chatbots are terrible at simple math and hallucinate is weirder than you thought

[email protected]

Someone put 69 to research and then to article. Nice trolling.

[email protected]

How I'd do it is basically

72 * (10+3)

(72 * 10) + (72 * 3)

(720) + (3*(70+2))

(720) + (210+6)

(720) + (216)

936

Basically I break the numbers apart into easier chunks and then add them together.

[email protected]

OK but the llm is evidently shit at math so its "non-standard" approach should still be adjusted

[email protected]

I am talking about the AI. It's already a computer. It shouldn't need to do anything other than calculate the equations. It doesn't have a brain, it doesn't think like a human, so it shouldn't need any special tools or ways to help it do math. It is a calculator, after all.

[email protected]

I wouldn't even attempt that in my head.
I can't keep track of things and then recall them later for the final result.

[email protected]

Anybody who claims they don't "think" before we even figure out completely how they work and even how human thoughts work are just spreading anti-AI sentiment beyond what is considered logical.

You should become a better example than an AI by only arguing based on facts rather than things you hallucinate if you want to prove your own position on this matter.

[email protected]

You're antropomorphising quite a bit there. It is not trying to be deceptive, it's building two mostly unrelated pieces of text and deciding the fuzzy logic is getting it the most likely valid response once and that the description of the algorithm is the most likely response to the other. As far as I can tell there's neither a reward for lying about the process nor any awareness of what the process was anywhere in this.

Still interesting (but unsurprising) that it's not getting there by doing actual maths, though.

[email protected]

I think it's odd in the sense that it's supposed to be software so it should already know what 36 plus 59 is in a picosecond, instead of doing mental arithmetics like we do

At least that's my takeaway

[email protected]

We also check to see if the word that popped into our heads actually rhymes by saying it out loud. Actual validation steps we can take is a bigger difference than being a little more robust.

We also have non-list based methods like breaking the word down into smaller chunks to try to build up hopefully more novel rhymes. I imagine professionals have even more tools, given the complexity of more modern rhyme schemes.

[email protected]

Yes, agreed. And calculators are essentially tabulators, and operate almost just like a skilled person using an abacus.

We shouldn't really be surprised because we designed these machines and programs based on our own human experiences and prior solutions to problems. It's still neat though.

? Offline

…Duh.

[email protected]

My favourite part of the day: commenting LLMentalist under AI articles.

[email protected]

It also doesn't help that the AI companies deliberately use language to make their models seem more human-like and cogent. Saying that the model e.g. "thinks" in "conceptual spaces" is misleading imo. It abuses our innate tendency to anthropomorphize, which I guess is very fitting for a company with that name.

On this point I can highly recommend this open access and even language-wise accessible article: https://link.springer.com/article/10.1007/s10676-024-09775-5 (the authors also appear on an episode of the Better Offline podcast)

[email protected]

Pen and paper maths I'm pretty decent at, but ask me to calculate anything in my head and it's anyone's guess if I remembered to carry the 1 or not. Ever since learning about aphantasia I'm wondering if the lack of being able to visually store values has something to do with it.

[email protected]

Times 5 and times 10 tables are really easy for me. So yeah, in my mind it's an easier comuptation.

That being said having a result of a little over a 1000 gives me an estimate for the magnitude of a number – it's around a thousand. It might be more or less but it's not far from there.

[email protected]

No it doesn't, multiplication and division always take precedence over addition and subtraction. You'd need parentheses to clarify what is in the divisor since that can be ambiguous with line notation.

[email protected]

"The planning thing in poems blew me away," says Batson. "Instead of at the very last minute trying to make the rhyme make sense, it knows where it’s going."

How is this surprising, like, at all? LLMs predict only a single token at a time for their output, but to get the best results, of course it makes absolute sense to internally think ahead, come up with the full sentence you're gonna say, and then just output the next token necessary to continue that sentence. It's going to re-do that process for every single token which wastes a lot of energy, but for the quality of the results this is the best approach you can take, and that's something I felt was kinda obvious these models must be doing on one level or another.

I'd be interested to see if there are massive potentials for efficiency improvements by making the model able to access and reuse the "thinking" they have already done for previous tokens

[email protected]

Ever since learning about aphantasia I’m wondering if the lack of being able to visually store values has something to do with it.

Here's some anecdotal evidence. Until I was 12 or 13, I could do absurdly complex arithmetical calculations in my head. My memory of it was of visualizing intermediate calculations as if they were on a screen in my head. I'd close my eyes to minimize distracting external stimuli. I'd get pocket money because my dad would get his friends to bet on whether I could correctly multiply two 7-digit phone numbers, and when I won, which I always did, he'd give the money to me. He had an old-school electromechanical calculator he'd use to check the results.

I was able to use a similar visualization technique to memorize long passages of music and text. That stayed with me post-puberty, though again at a lesser extent.

Once puberty kicked in, my ability to visualize declined significantly, though I also learned some mental arithmetics tricks that I still use now. I was able to get an MS in mathematics without much effort, since that relied on higher-level reasoning and not all that much on powerful memory or visualization.

So I think your comment about aphantasia is at least directionally correct, as applied to people. But there's little reason to assume LLMs would do things the same way a human mind does, though both might operate under similar information-theoretic constraints.

[email protected]

The research paper looks well written but I couldn’t find any information on if this paper is going to be published in a reputable journal and peer reviewed. I have little faith in private businesses who profit from AI providing an unbiased view of how AI works. I think the first question I’d like answered is did Anthropic’s marketing department review the paper and did they offer any corrections or feedback? We’ve all heard the stories about the tobacco industry paying for papers to be written about the benefits of smoking and refuting health concerns.

[email protected]

Memory can improve with training, and it's useful in a large number of contexts. My major beef with rote memorization in schools is that it's usually made to be excruciatingly boring. I'd say that's the bigger problem.

agnos.is Forums

Anthropic has developed an AI 'brain scanner' to understand how LLMs work and it turns out the reason why chatbots are terrible at simple math and hallucinate is weirder than you thought