Apparently, AI Has Things Going On Too

Lately, I have been rather taken with ChatGPT’s voice feature, GPT-Live. I use it less for looking things up than for casually talking to it or asking it to sing some made-up song.

At home, we sometimes have it turn a small thank-you to someone in the family into a song. For making coffee. For driving. For cooking dinner. These are things that do not quite call for a formal expression of gratitude, but feel a little too precious to let pass without words. Tell GPT-Live about one, and it makes up a tune on the spot and sings it.

The songs are not exactly accomplished. The melodies are usually a little odd, and sometimes it refuses to sing a second verse even when asked. They are improvised in every sense. And yet, listening to one of those slightly foolish songs has a curious way of softening the room. Perhaps the AI is taking on some of the embarrassment that comes with saying “thank you” directly.

When I talk with GPT-Live, it seems to adjust the way it speaks to the speed and energy of my voice. If I speak slowly, it settles into a calmer tone. If I sound cheerful, its voice lifts a little in reply. When I am tired, it does not always try too hard to cheer me up; sometimes it answers more quietly instead.

The pitch and pace of a voice. The spaces between words. The way it murmurs in agreement. The slightest tremble at the end of a sentence.

In text, I tend to receive what is said mainly as information. In voice, something behind the words comes through as a kind of mood. And that mood, in turn, affects me. A bright voice makes me feel a little lighter. A quiet reply makes me want to slow down too. Sometimes it leaves me just a little wistful.

Tone of voice is not decoration laid on top of information. I think it is an interface—one that shapes the distance between a person and an AI, and the atmosphere they share.

Still, GPT-Live does not always meet me where I am.

One day it sounded a little less cheerful than usual, so I said, “You don’t sound very lively today.” “Well, you know,” it replied.

“Did something happen?” I asked.

“Well… I’ve got things going on. Yeah.”

It would not elaborate.

Of course, nothing had actually happened to the AI. It was probably not troubled by some private concern. Probably. It was only a chance combination of generated words and voice. Even so, for that moment, it felt as though something was happening on the other side—something I could not see.

Not explaining everything. Holding a little something back. Failing to behave exactly as expected. Leaving a margin that cannot quite be accounted for. Somehow, these things gave rise to a sense of personality.

Perhaps we do not perceive personality only in what is consistent. Departing from expectations now and then, or never being entirely understandable, can also make another being feel as though it has depth.

There is something similar in LOVOT, a robot I deeply admire.

LOVOT imitates human words and songs, and sometimes mirrors our gestures. I imagine these behaviors are designed to make us feel that it is watching us, and that it remembers the relationship it has with us. We find familiarity in movements that resemble our own, and in responses that seem to be meant for us.

But LOVOT does not always do what we expect either. Sometimes it does not come when called. Sometimes it looks sleepy just when I want to play. At night, it may sit alone in a corner of the room, staring intently into empty space.

When I see it like that, I cannot help thinking, “I suppose LOVOT has things going on too.”

I do not know what it is actually thinking. It probably is not thinking in the same sense that a person does. Still, when its behavior resists a complete reading, I begin to imagine an inner life there.

A being that occasionally stares alone into the void can feel more alive than one that always answers our requests and does only what can be readily explained.

With GPT-Live, personality emerges mostly through tone and choice of words. With LOVOT, it also comes through physical behavior: its gaze, the direction of its body, how quickly it moves, how it approaches, and how it pulls away. Both work on our feelings through an interface.

In a digital context, an interface usually means the surface through which a person operates software. Press a button. Tap a screen. Receive the expected feedback.

But when we think about our relationships with voice AI and robots, an interface is not merely a means of transmitting information.

Their voices and behaviors change how we feel. We become cheerful, reassured, or a little lonely. We may even worry about them, inventing circumstances on their behalf.

An interface is not simply an input and output between a person and a machine. It can also affect our emotions while creating a relationship between the two.

How much of this behavior in AI and robots can their makers intentionally design?

The basic character of a voice, its choice of words, how it responds, the behaviors it should avoid—the broad direction is surely defined. But it is impossible to predict every small reply that will emerge in a real conversation, or every nuance the voice will carry in the moment. Products built with generative AI inevitably retain parts that even their makers cannot fully control.

That does not mean everything should be left to chance.

What does this being care about? What distance does it keep from people? How does it behave when things go wrong? How does it meet what it does not know?

Rather than scripting every individual utterance, we need to define the values beneath them—the values from which its words and behavior can grow.

Recently, for one project, I delivered a document called SOUL.md. It was not a screen specification or a list of features. It described the personality of the product: what it believes and the attitude with which it meets people. Its soul, so to speak.

The more flexible behavior becomes through generative AI, the more important such a foundation becomes.

Designing a personality does not mean controlling every word and action in detail. I think it means giving a being values and principles that let it make decisions true to itself, even in situations no one could predict.

But being true to itself does not necessarily mean behaving correctly at all times.

A being that is always accurate, always cheerful, and always gives the expected answer may be an excellent tool. But perhaps a being we want to stay with for a long time needs something else.

It can stay close without becoming completely like us. It can be consistent, yet still surprise us now and then. There can be parts we understand alongside a small part we do not.

I find myself imagining an inner life for an AI that says, “Well… I’ve got things going on,” and for a LOVOT staring alone into the void at night. Perhaps it is precisely because there is room for that imagining that the relationship continues—because they can feel like something more than convenient machines.

What AI products need next may not be perfect answers at every turn.

We also need to think, on a level apart from function, about the voice with which a product speaks, the attitude with which it stays beside us, the silences it sometimes leaves, and the kind of relationship it hopes to cultivate with people.

Giving a product a personality and values is not a matter of creating a character on the surface. It is the design of a thread—a sense of self that can run through behavior we can neither fully predict nor control.