Tavus has rolled out Griffin, a unified video-to-video system that combines recognition, generation, and logic, bypassing the traditional model cascade. Forget about cumbersome chains of separately operating transcription models, language engines, and speech synthesizers. The new architecture makes decisions several times per second without waiting for your interlocutor to finish their passionate speech.

In practice, this looks frighteningly lifelike: the system reacts to events on screen, including screen sharing, nods along during a monologue, and instantly falls silent if you interrupt it. To launch this digital twin, developers need only a single photograph and ten seconds of your audio. According to Tavus's own data, 48% of closed-test participants genuinely believed they were communicating with a live human. Against the backdrop of the previous technological stack, where this metric barely reached 2%, the progress is as impressive as it is alarming.

The Griffin-Lite version is currently in closed preview, but its commercial prospects are already causing serious headaches for security specialists. This level of realism greenlights a new generation of real-time deepfake fraud. Business leaders should review video call verification protocols right now, because trusting your own eyes and ears on screens is no longer an option.

Artificial IntelligenceGenerative AICybersecurity