← All stories

Multimodal AI - a world beyond LLMs

19 August 2026
Multimodal AI - a world beyond LLMs

One hour of recorded speech takes six to ten hours to label by hand, and nobody can do it for more than four hours a day before the errors creep in. Teodora Vuković did it through her entire PhD. The sentences played on a loop in her dreams for six years.

Teodora is a linguist by training who leads the multimodal technology group at the University of Zurich and founded MOSAIC. She came to AI from an observation every linguist eventually makes: words alone never carried the meaning. We speak with our faces, our hands, our tone, and for most of computing history machines understood none of it.

Xaver Walser, who was in the room, reached for cars to explain what that changes. The first automobiles were carriages with an engine bolted in where the horse used to be. The body only found its own shape once designers built around the engine instead of around the animal. Large language models are the engine we have been designing around for three years, which is why almost everything built on them is a text box: a chat window, a prompt, a document. Multimodal AI is a different engine. What gets built on top of it will not look like a chat window.

Teodora Vuković presenting to the Ship26 crew at Schifflände 26, a slide behind her asking the room what they think multimodal AI is

The paradox she unpacked for the room is that everything a human does without thinking had to be spelled out for the machine, label by label. You register whether a speaker is relaxed or nervous without any awareness of doing it. A machine needed millions of examples, each one described by a person.

That gap explains why the field moved so slowly, and in a specific order. Text got there first because text describes itself: it arrives pre-segmented into words and sentences, and language carries its own labels. Speech has no natural units and breaks on real-world noise, an air conditioner humming, a boat passing outside, both of which were happening in the room as she said it. Images are a grid of pixels, and their labelling bottleneck was so severe the industry outsourced it to all of us, one CAPTCHA at a time. Video sits at the top: a machine still cannot reliably tell where one gesture ends and the next begins. As Xaver put it afterwards, our most tedious manual task today is almost certainly our next product tomorrow.

Then the finding that surprised her own team. Strip a face of its appearance entirely, keep nothing but how the points move, and the movement identifies a person with 90 to 95 percent accuracy. Your face gives you away, even without your face.

Sit with that for a second, because a whole practice rests on the opposite assumption. Research interviews get published with the face blurred and the background blacked out. Anonymisation, as we have understood it for decades, means removing what a person looks like. It was designed against the old engine. And a blurry figure in the far background of a video used to be nobody: data that cannot be processed today will be processable tomorrow.

The room stress-tested all of it. After the live demo someone asked, point blank, what problem this solves. Someone else called it a polygraph on steroids. Teodora took the fire with humour and stayed grounded on the limits: models still read unfamiliar cultures poorly, and a smile in the data says nothing about whether it is joy or someone smiling through pain.

And on how these systems are actually built, who better to answer than someone who works on them. In the room with us sat somebody working on Google Gemini, who opened up the engineering underneath: text, sound and image all becoming tokens in one shared space.

The talk officially ran out of time. Nobody left. The conversations kept going for another thirty minutes, from coaching and therapy platforms circling exactly this technology, more than one person in the room is already building or sketching in that space, to the harder question of how much of it should sit with big tech. Thank you, Teodora, for handing the room a field most of us could not have described that morning, and for taking an hour of pushback without once retreating into jargon.

The last word came from somewhere in the back: this will go the way of personalised medicine. Years of your own data, feeding an agent that knows what your smile means, not what smiles mean on average. Teodora adopted it on the spot as her closing line. A new engine, and the room already sketching the body around it.

Wednesdays at Ship26 tend to run like this. What's coming up next is on our events page.

Our Values

S(Stewardship)Build for the generation to come.
H(Humanity)People before pipelines.
I(Imagination)See what isn't there yet, treat this seeing as part of the work.
P(Possibility)If we can imagine a better future, we can build it.