Multimodal AI - a world beyond LLMs
One hour of recorded speech takes six to ten hours to label by hand, and nobody can do it for more than four hours a day before the errors creep in. Teodora Vuković did it through her entire PhD. The sentences played on a loop in her dreams for six years.
Teodora is a linguist by training who leads the multimodal technology group at the University of Zurich and founded MOSAIC. She came to AI from an observation every linguist eventually makes: words alone never carried the meaning. We speak with our faces, our hands, our tone - and for most of computing history, machines understood none of it.
Multimodal AI may not be something many of us were familiar with - until this Wednesday. Thanks to Teodora's talk, the room understands the field a whole lot better. Large language models are great at interacting in written text; multimodal AI takes this several steps further - and meanwhile, machines can identify you by nothing but how you move.

The paradox she unpacked for the room is that everything a human does without thinking had to be spelled out for the machine, label by label. You register whether a speaker is relaxed or nervous without any awareness of doing it. A machine needed millions of examples, each one described by a person.
That gap explains why the field moved so slowly - and in a specific order. Text got there first because text describes itself: it arrives pre-segmented into words and sentences, and language contains its own labels. Speech has no natural units and breaks on real-world noise - an air conditioner humming, a boat passing outside, both of which were happening in the room as she said it. Images are a grid of pixels, and their labeling bottleneck was so severe the industry outsourced it to all of us, one CAPTCHA at a time. Video sits at the top: a machine still can't reliably tell where one gesture ends and the next begins.
Then the finding that surprised her own team. Strip a face of its appearance entirely - keep nothing but how the points move - and the movement identifies a person with 90 to 95 percent accuracy. Your face gives you away, even without your face.
The implication followed quietly: a blurry figure in the far background of a video used to be nobody. Data that can't be processed today will be processable tomorrow.
The room stress-tested all of it. After the live demo someone asked, point blank, what problem this solves. Someone else called it a polygraph on steroids. Teodora took the fire with humor and stayed grounded on the limits: models still read unfamiliar cultures poorly, and a smile in the data says nothing about whether it is joy or someone smiling through pain.
And on how these systems are actually built - who would be better suited to answer than someone who works on them? In the room with us sat somebody working on Google Gemini, who opened up the engineering underneath: text, sound and image all becoming tokens in one shared space.
The talk officially ran out of time, but the conversations kept going for another 30 minutes, from coaching and therapy platforms circling exactly this technology - more than one person in the room is already building or sketching in that space - to the harder question of how much of it should sit with big tech.
The last word came from somewhere in the back of the room: this will go the way of personalized medicine. Years of your own data, feeding an agent that knows what your smile means - not what smiles mean on average. Teodora adopted it on the spot as her closing line.
Wednesdays at Ship26 tend to run like this. What's coming up next is on our events page.