Skip to main content

SBS AI Series | Q&A with Sandiway Fong: Human Language and AI

Monday
Image
Sandiway Fong on the left with a graphic of a computer generated face that is speaking on the right

Artificial intelligence is changing rapidly, but understanding its impact requires looking beyond the technology itself. Across the College of Social and Behavioral Sciences, our faculty are examining AI through the lenses of ethics, language, culture, human behavior, policy and society — and asking fundamental questions about what it means to be human in an age of intelligent machines. As part of our exploration of AI and its many facets, this series features in-depth conversations with SBS scholars whose expertise brings distinct perspectives to this rapidly changing field. These conversations go beyond the headlines to explore the questions, possibilities and consequences of AI — and what they can tell us about ourselves and the world we are building.

In this third of three Q&As, Sandiway Fong, director of the Human Language Technology Program in the Department of Linguistics, brings his expertise in computational linguistics to the conversation about AI. He talks about what AI can — and cannot — tell us about human language, whether AI truly understands language and what its quickly evolving advances show us about the difference between human thought and machine learning. 

 

AI can write emails, stories, even jokes that sound surprisingly human. Does that mean it actually understands, or is “learning” language?

This is a controversial topic. Suppose we tackle it this way: If a parrot repeats human speech, does it understand language? I assume we agree it does not. To the point, memorization (and the ability to vocalize) does not constitute understanding. Similarly, John Searle’s famous Chinese Room thought experiment tells us that purely mechanical manipulation does not mean we understand Chinese. Then what is understanding? To a first approximation, understanding requires interpretation and thinking. Because AI models are notoriously opaque, it is currently impossible to pinpoint where “thinking” occurs within that vast apparatus of interacting floating-point matrices and tokenized language units. But because it makes context-appropriate and lucid replies, AI proponents tell us it must be understanding language.

To address the second point, AI systems require vast amounts of human-generated language data to perform adequately. The scale of such machine learning systems is staggering both in terms of compute power and data. By comparison, humans are a minor miracle of nature. We manage to somehow learn language with a slow biological brain that takes only 20W of power, and we converge on language on surprisingly little exposure to primary linguistic data and explicit instruction. Moreover, we are capable of beautiful and surprising invention when it comes to language. We can express new thoughts, say sentences and use new words that the hearer has never heard before, and yet the hearer can understand our inner thoughts and intentions. 

Today, AI systems output what is known as AI slop. It’s a problem as AI systems greedily reach for more training data to try to scale up further. AI systems try to filter out AI-generated slop and preferentially use human-generated language in training. You may have heard recently that AI companies have resorted to buying rare books and perhaps destroying them in the process of taking them apart for scanning.

 

As someone who studies both linguistics and AI, has working with language models changed the way you think about human language?

As a regular computer scientist by training (Ph.D. MIT AI Laboratory) and a computational linguist (someone who builds computer models of language), I can see that these two fields have radically different goals. Linguists want to understand how language works in the human brain, why language is this way and not some other (random) way. Also, how we can acquire language and make inventive use of it given the so-called Poverty of Stimulus? 

Like much of natural science, a deep explanation, couched in terms of simple biological mechanisms, is the goal. And, of course, we need an explanation of how evolution managed to give us these mechanisms (and presumably, as far as we can tell, no other living creatures). Noam Chomsky, a long-time professor at MIT and a faculty member at U of A, has perhaps been the most influential thinker in how we should go about constructing such theories scientifically.

On the other hand, computer scientists are interested in general machine learning methods, rather than mechanisms specific to language. AI systems are an illustration of their success in applying these methods. A Nobel Prize was awarded to computer scientist Geoffrey Hinton in 2024 for foundational work in how to train AI systems. But AI systems do not operate under the austere conditions of the human brain, i.e., 20W of power and limited exposure to community language. And so brute-force methods cannot provide an explanation for human language. There is much more to be said, but in this respect, Hinton is clearly mistaken and Chomsky is correct. 

You may recall that machine supremacy in chess was established back in 1997 when a supercomputer beat the human world champion. Note there are no machine vs. world champion chess matches these days. Did it teach us anything about human thought? Did that knowledge transfer to other fields (as was promised by IBM Research)? The simple answer is nothing came of it. Chomsky said at the time, I’m paraphrasing here, that all it did was take the fun out of playing chess.

 

Has AI confirmed anything linguists have long believed about language — or challenged those ideas?

Although AI does not directly address the scientific goals of linguistics as stated above, and therefore cannot speak to those goals, we can be sure it has had major impact on linguistics. 

Just one example, Hinton claims both that humans must use the same (successful) mechanisms as AI systems, i.e., essentially that the problem of language has been solved, and that 70 years of modern linguistic theory has largely failed to provide (mechanizable) alternative mechanisms. This has obvious research funding consequences. However, the recent recognition of AI slop is perhaps a sign that AI-generated language is not quite the same as human language. Also recently, the topic of plagiarism and AI-generated prose has made the news, e.g., the tragic case of Cambridge University professor Jason Arday. 

There are AI programs that purport to distinguish between human and AI-generated language. Obviously, people believe it is (still) possible to distinguish between the two. If so, there is a visible gap between them. Professors today are faced with the dilemma of determining whether submitted homework and essays are the work of students or AI or (increasingly likely) a mix of both.

 

When it comes to language, where has AI made the most progress and where do you think it still has the furthest to go?

Right now, we can be fairly sure AI systems have harvested every readily accessible part of human language, including the case of rare books mentioned earlier. AI systems generate lucid and grammatical output for many human languages, can recite passages word-for-word and do a valuable job of translation, captioning and summarization (though one can readily find humorous failures). All previous attempts to mechanize these useful tools have failed to reach usefulness. 

"This is a spectacular triumph of advances in computer hardware engineering and machine learning and rightly should be celebrated. However, the next step should be thought and that perhaps is where it has the furthest to go."

One way to think about the question is to ask, what is human language really for and why did it come about? There is a rather famous historical controversy, pitting those who believe language naturally evolved for communication against those who believe language is actually for internal thought, e.g., Chomsky. There is much to be said here that we don’t have time or space for, but suppose for a moment that Chomsky is right. Homo sapiens alone have that rich, inventive ability with language that provides for internal thought, and this is why there exists the explosion of symbolic artifacts in the fossil record that coincides with our emergence as a species (after 4 billion years of evolution). 

Well, if language gives us our species-specific ability for innovative thought, should not AI also possess that ability, assuming it really possess language? This, to me, is a most fascinating question for the future.

 

##