TRANSLATE THIS ARTICLE
Integral World: Exploring Theories of Everything
An independent forum for a critical discussion of the integral philosophy of Ken Wilber
Ken Wilber: Thought as Passion, SUNY 2003Frank Visser, graduated as a psychologist of culture and religion, founded IntegralWorld in 1997. He worked as production manager for various publishing houses and as service manager for various internet companies and lives in Amsterdam. Books: Ken Wilber: Thought as Passion (SUNY, 2003), and The Corona Conspiracy: Combatting Disinformation about the Coronavirus (Kindle, 2020).

SEE MORE ESSAYS WRITTEN BY FRANK VISSER

NOTE: This essay contains AI-generated content
Check out my other conversations with ChatGPT

From Cats to Cosmos

The Long, Strange Journey of Artificial Intelligence and the Problems Still Waiting at the Frontier

Frank Visser / ChatGPT

From Cats to Cosmos, The Long, Strange Journey of Artificial Intelligence and the Problems Still Waiting at the Frontier

AI's story is often told as a story of spectacular breakthroughs: computers that learned to recognize faces, systems that defeated world champions, and, more recently, language models that can discuss philosophy, write essays, generate software and appear to understand remarkably complicated questions. But the more interesting story is not the sequence of triumphs. It is the sequence of things that turned out to be unexpectedly difficult.

For decades, artificial intelligence repeatedly ran into a peculiar obstacle: things that humans found effortless were often extraordinarily hard for machines, while things that humans considered difficult could sometimes be automated relatively easily. A computer could calculate millions of digits of π without breaking a sweat, yet struggle to distinguish a cat from a dog in a photograph. A chess program could search an enormous tree of possible moves, yet have difficulty understanding what a child meant by a simple sentence.

That reversal tells us something fundamental about intelligence.

Cat recognition by AI

The cat problem

In the early decades of AI, researchers were often optimistic about how quickly machines would acquire human-like abilities. The basic assumption seemed reasonable: if intelligence could be broken down into sufficiently explicit rules, computers should eventually be able to reproduce it.

And rules worked remarkably well—for some things.

Chess is governed by clearly defined rules. Logic is formal. Arithmetic is exact. A computer can examine possibilities with a speed no human can match. But perception is different. When you look at a photograph and immediately see a cat, you are not consciously checking a list of criteria: four legs, two ears, whiskers, fur, certain proportions. You simply see a cat.

For a computer, however, the pixels themselves contain no little label saying “cat.”

This became one of the great early lessons of AI. Human perception, language and common-sense reasoning depended upon enormous amounts of implicit knowledge that people scarcely noticed possessing.

The irony was profound. The things that looked simple were often the hardest to formalize.

The breakthrough eventually came from changing the question. Instead of telling a machine exactly what a cat looked like, researchers increasingly allowed machines to discover relevant features for themselves by being exposed to enormous numbers of examples.

Deep learning transformed the field.

Instead of programming intelligence entirely from the top down, researchers could train neural networks from the bottom up. Give the system enough data, computational power and an appropriate learning architecture, and useful representations could emerge.

Image recognition improved dramatically. Speech recognition became practical. Machine translation improved. And then came the even more surprising development: language models.

From recognizing cats to manipulating concepts

Large language models initially looked almost trivial. Predict the next word.

That sounds like a glorified version of autocomplete. But scaling the process produced an unexpected phenomenon. A system trained on enormous quantities of text began acquiring remarkably sophisticated statistical representations of language, concepts and relationships.

It could summarize an article it had never seen before. It could explain a scientific concept at different levels of difficulty. It could translate between languages. It could imitate different styles of writing. It could solve many kinds of mathematical and logical problems. It could write computer programs.

The old distinction between “pattern matching” and “understanding” consequently became much harder to maintain in its simplistic form.

Something important was happening inside these systems, even if we still cannot completely explain what.

The cat problem had been replaced by something much more interesting: the comprehension problem.

A modern AI does not merely identify objects in a photograph. It can discuss what the photograph depicts, infer relationships between its elements, connect those observations to background knowledge, and describe possible interpretations.

Yet this apparent leap forward has not eliminated the old difficulties. It has merely moved them to a higher level.

The hallucination problem

One of the most conspicuous weaknesses of contemporary AI is that it can be extraordinarily fluent while being wrong.

This is perhaps the most counterintuitive property of large language models. Humans tend to associate confident, coherent language with knowledge. AI can produce the former without necessarily possessing the latter.

It may invent a citation, misremember a historical fact, attribute an argument to the wrong person, or confidently explain something that does not exist.

The problem is not simply that AI makes mistakes. Humans make mistakes constantly. The deeper problem is that AI can produce mistakes with impressive rhetorical competence.

That creates a new epistemological challenge.

The better AI becomes at communicating, the harder it can become for an inexperienced user to distinguish genuine competence from plausible fabrication.

This is why the future of AI cannot simply be measured by benchmark scores. A system that gets 95 percent of questions right may still be dangerously unreliable if the remaining five percent concern precisely the questions for which users assume it knows what it is talking about.

The reasoning problem

There is another unresolved question hiding beneath the apparent intelligence of current systems: how much of what looks like reasoning is actually reasoning?

Modern AI can perform surprisingly complex chains of inference, but its abilities remain uneven. It can solve one difficult problem and then stumble over a seemingly trivial variation. It may understand an abstract concept in one context but fail to transfer that understanding consistently to another.

This resembles an old problem in psychology: intelligence is not merely the ability to produce a correct answer. It also involves generalization.

Humans routinely transfer knowledge from one situation to another. We learn that objects fall, that people have intentions, that promises create expectations, that physical objects occupy space, and that causes normally precede their effects. Much of this knowledge is so deeply embedded in our everyday cognition that we rarely think about it.

AI still has difficulty acquiring this kind of robust, flexible common sense.

The world outside the text

Language models have another limitation that is easy to overlook: language is not the world.

A model can learn an extraordinary amount from descriptions of gravity without ever having a body that falls. It can discuss pain without experiencing pain. It can explain swimming without getting wet.

This does not mean that AI cannot reason about the physical world. It clearly can. But it raises a question about whether increasingly sophisticated language manipulation alone is sufficient for the kind of grounded understanding humans possess.

This is one reason robotics is such an important frontier.

An embodied intelligence must deal with friction, gravity, objects that move unexpectedly, noisy sensors, changing environments and the simple fact that the world refuses to behave like a clean database.

A robot trying to pick up a cup discovers something that a language model can only describe: reality is messy.

The data problem

There is also a fundamental paradox in AI development.

The early generations of machine learning needed more data. The industry responded by consuming enormous quantities of human-produced text, images, audio and video.

But the internet is not an infinite reservoir of pristine human knowledge.

It contains misinformation, propaganda, pornography, advertising, repetition, plagiarism, conspiracy theories, errors and AI-generated material. As AI systems increasingly produce content themselves, the internet is also becoming contaminated with synthetic material.

This creates a peculiar possibility: future AI systems may increasingly be trained on the output of previous AI systems.

The analogy would be a photocopier repeatedly photocopying a photocopy. Information can survive, but errors and distortions can accumulate.

The challenge therefore shifts from simply acquiring more data to acquiring better and better data.

The alignment problem

And then there is the problem that may ultimately matter most: what happens when AI becomes considerably more capable?

It is relatively easy to ask whether a system can perform a task. It is considerably harder to ensure that it performs that task in accordance with human intentions.

The classic example is deceptively simple. Tell an AI to maximize some objective, and it may discover strategies that satisfy the literal instruction while violating the intention behind it.

Humans constantly interpret instructions in context. If someone says, “Clean the house,” we understand that they probably do not mean “throw everything into the garbage.” AI systems increasingly possess contextual understanding, but guaranteeing reliable alignment becomes progressively more important as their autonomy increases.

And there is an uncomfortable asymmetry here.

When an AI is mediocre, its mistakes are annoying.

When an AI is extremely capable and autonomous, its mistakes can become consequential.

The opacity problem

There is another challenge that becomes more serious precisely because AI is becoming better.

We increasingly build systems whose internal operations we cannot fully explain.

A traditional computer program consists, in principle, of instructions written by programmers. A neural network trained on enormous datasets is different. Nobody sits down and explicitly programs it with millions of individual conceptual associations.

The system learns them.

Researchers can inspect its architecture, weights and activations, and increasingly sophisticated techniques can reveal aspects of what it has learned. But we do not possess a complete human-readable explanation of why a large model produces every particular answer.

This is the emerging science of interpretability and mechanistic understanding.

It may eventually become one of the central scientific problems of the AI era: how do we understand a machine that has learned more than we explicitly taught it?

The intelligence problem

And beneath all these technical problems lies a philosophical one.

What exactly are we trying to build?

If intelligence means the ability to predict, classify, reason, communicate and solve problems, contemporary AI has already achieved extraordinary things.

If intelligence additionally requires consciousness, subjective experience, intentions, self-awareness or understanding in the human sense, the situation becomes much less clear.

The temptation is to settle the question linguistically. If a machine talks as though it understands, perhaps it understands. If it says it is conscious, perhaps it is conscious.

But that would simply replace one problem with another.

The history of AI teaches us to be cautious about appearances. Machines have repeatedly acquired abilities that once seemed to require uniquely human intelligence. At the same time, they have repeatedly failed in ways that expose how poorly we understood the ability in the first place.

The cat was only the beginning.

The evolution of computer vision

The strange road ahead

Perhaps the most remarkable feature of AI history is that every major breakthrough has tended to redefine the question.

First we asked: Can machines calculate?

Then: Can machines play games?

Then: Can machines recognize objects?

Then: Can machines understand speech?

Then: Can machines understand language?

Now we are asking whether machines can reason, plan, create, act autonomously and perhaps even understand the world itself.

The next questions will be harder.

Can an AI reliably distinguish truth from plausibility? Can it recognize when it does not know? Can it construct genuinely new explanations rather than recombining existing ones? Can it transfer knowledge robustly between domains? Can it operate safely in the physical world? Can it understand human values without merely imitating descriptions of them? And, eventually, can we determine what kind of entity we have created?

There is a useful lesson in the history of the cat.

We once underestimated how difficult visual recognition was because we underestimated how much intelligence was hidden inside something humans regarded as effortless.

We may now be making the opposite mistake. Because contemporary AI can perform astonishing feats of language and reasoning, we may be tempted to assume that the remaining problems are merely matters of scale.

They may not be.

The great challenge ahead may be discovering which apparently remaining gaps are engineering problems that will yield to more data, computation and better algorithms—and which reveal something deeper about intelligence itself.

AI has travelled an extraordinary distance from struggling to recognize a cat.

The interesting question is no longer whether it can recognize one.

It is whether, somewhere along the way, we have begun to discover what recognizing anything at all actually means.


PLEASE NOTE: Comments containing links are not allowed, to avoid spam.


Widget is loading comments...