The car could understand me.

That was supposed to be the revolution.

Twenty-five years ago, talking to a car was a little like talking to a particularly stupid receptionist. You had to use the right words, in roughly the right order, and preferably without sounding too human.

“Open the window.”
That worked.
“Could you open the window?”
Maybe.
“I’m getting a little warm in here. Mind opening my window?”
The car would stare back at you with the electronic equivalent of a blank face.
“I can’t understand you.”

We got better at this. Much better. Today, an LLM can understand what you mean even when you don’t say it the way the engineers expected. You can combine requests, refer to something you said thirty seconds ago, correct yourself, and leave things implicit. You can talk like a human. And yet, in many cases, you’re still talking to a command-and-control system. The command has simply learned better manners.

The old systems were built around commands: an intent, some parameters, an action, and a response. The user said something that was mapped to a function. The function was executed. The system said something back. LLMs have made the first part dramatically better. Instead of learning the vocabulary of the machine, we can increasingly use our own.

“Open the driver’s window.”
“I’m getting stuffy.”
“Could you let some air in?”
“Can you crack my window?”

These can all essentially mean the same thing. That is a huge improvement. But easier commands do not automatically make an interaction conversational.

There is a difference between understanding natural language and having a natural interaction.

Imagine a human passenger sitting next to you.

You say, “I’m getting cold.”
They might say, “Want me to turn the heat up?”
You say, “A little.”
They know what “a little” means because there is a shared context. They turn the temperature from 20 to 21 degrees.
You say, “No, that’s too much.”
They stop.
You say, “Actually, just leave it.”
The conversation changed direction.

Nobody has to press a button labeled CANCEL_CURRENT_INTENT. Nobody has to restart the dialogue or wait for a sentence to finish before speaking again. The important word here is interaction.

A modern LLM can understand a remarkably complicated command:

“Set the temperature to 21, turn on the seat heating for the driver, call John, and when he answers, tell him I’ll be about ten minutes late.”

Twenty-five years ago, that would have been science fiction. Today, it is largely a language-understanding problem.

But suppose the car answers:

“Okay. Setting the temperature to 21 degrees, activating the driver’s seat heating, calling John, and…”

“Actually, don’t call him.”

Now what? A genuinely conversational system simply stops. In a command-and-control system, interruption can still be a special case. The conversation is really a sequence of transactions wearing a conversational overcoat.

This becomes even more obvious when the system has to make a decision with you.

“Take me to the office.”
“There is heavy traffic on your usual route. I can take an alternative route that is approximately five minutes faster, but it includes a toll.”
“Fine.”
“Okay.”

Here the car isn’t merely translating language into an API call. The user has expressed a goal. The system has identified a problem, generated an alternative, and asked for a decision. The difference is subtle until you experience it.

Today’s LLM-based voice interfaces can make this distinction easy to miss. Natural pauses, incomplete sentences, remembered context, and even small talk can make a system sound conversational. But a parrot can sound conversational. The interesting question is what happens when the conversation stops following the script.

Can you interrupt it? Change your mind? Pick up something said three turns ago? Ask a question instead of executing a command? Switch languages halfway through?

You’re driving in Germany and talking to the car in English.
“Take me home.”
The car asks a clarifying question.
You answer in German.
„Nein, ich meine die andere Adresse.“

A human doesn’t need a language-switching ceremony. The conversation just continues.

Modern real-time agentic systems are beginning to make this kind of interaction feel natural. Frameworks and APIs such as OpenAI’s Realtime API make it possible to build around persistent, low-latency interaction rather than treating every utterance as a separate request. They can listen while speaking, handle interruptions, maintain conversational state, and decide what to do next.

The technology, in other words, increasingly allows the conversation itself to become the interaction model. Whether automotive systems actually adopt that model is another question.

The automotive industry is beginning to move in this direction. Volkswagen is explicitly talking about “Agentic AI,” Mercedes-Benz is introducing Google’s Automotive AI Agent into MBUX, and BMW is integrating Amazon Alexa+ into its Intelligent Personal Assistant. These developments point beyond simply putting an LLM behind a voice-command interface, toward systems that can maintain context, handle multi-turn interaction, and act on broader user goals.

But the transition is still underway. Much of what we call conversational automotive AI remains fundamentally task-oriented.

Humans, meanwhile, are remarkably bad at behaving like APIs. We interrupt each other. We change our minds. We leave sentences unfinished. We switch languages. We correct ourselves halfway through a thought. We say things that make sense only because of what happened thirty seconds earlier.

For decades, automotive voice control has tried to make humans behave more like software. LLMs have finally made software better at dealing with humans. We may still, however, have built a much better command interface rather than a true conversational partner. The command grammar is gone. The command-and-control paradigm isn’t.

The first generation of voice assistants asked, “What command did you say?”

The LLM generation asks, “What did you mean?”

The next step is broader: “What are we trying to accomplish?”

That moves the car from executing instructions toward participating in an interaction. Understanding the sentence is no longer enough. The system has to understand that the sentence is part of a conversation.

And the funny thing is that we may have spent twenty-five years teaching cars to understand what we say, only to discover that the real breakthrough is teaching them how to listen.