AI video just took a startling leap in realism. Are we doomed?
2 day ago / Read about 51 minute
Source:ArsTechnica
Google's Veo 3 delivers AI videos of realistic people with sound and music. We put it to the test.


Credit: Google

Last week, Google introduced Veo 3, its newest video generation model that can create 8-second clips with synchronized sound effects and audio dialog—a first for the company's AI tools. The model, which generates videos at 720p resolution (based on text descriptions called "prompts" or still image inputs), represents what may be the most capable consumer video generator to date, bringing video synthesis close to a point where it is becoming very difficult to distinguish between "authentic" and AI-generated media.

Google also launched Flow, an online AI filmmaking tool that combines Veo 3 with the company's Imagen 4 image generator and Gemini language model, allowing creators to describe scenes in natural language and manage characters, locations, and visual styles in a web interface.

An AI-generated video from Veo 3: "ASMR scene of a woman whispering "Moonshark" into a microphone while shaking a tambourine"

An AI-generated video from Veo 3: "ASMR scene of a woman whispering "Moonshark" into a microphone while shaking a tambourine"

Both tools are now available to US subscribers of Google AI Ultra, a plan that costs $250 a month and comes with 12,500 credits. Veo 3 videos cost 150 credits per generation, allowing 83 videos on that plan before you run out. Extra credits are available for the price of 1 cent per credit in blocks of $25, $50, or $200. That comes out to about $1.50 per video generation. But is the price worth it? We ran some tests with various prompts to see what this technology is truly capable of.

How does Veo work?

Like other modern video generation models, Veo 3 is built on diffusion technology—the same approach that powers image generators like Stable Diffusion and Flux. The training process works by taking real videos and progressively adding noise to them until they become pure static, then teaching a neural network to reverse this process step by step. During generation, Veo 3 starts with random noise and a text prompt, then iteratively refines that noise into a coherent video that matches the description.

AI-generated video from Veo 3: "An old professor in front of a class says, 'Without a firm historical context, we are looking at the dawn of a new era of civilization: post-history.'"

AI-generated video from Veo 3: "An old professor in front of a class says, 'Without a firm historical context, we are looking at the dawn of a new era of civilization: post-history.'"

DeepMind won't say exactly where it sourced the content to train Veo 3, but YouTube is a strong possibility. Google owns YouTube, and DeepMind previously told TechCrunch that Google models like Veo "may" be trained on some YouTube material.

It's important to note that Veo 3 is a system composed of a series of AI models, including a large language model (LLM) to interpret user prompts to assist with detailed video creation, a video diffusion model to create the video, and an audio generation model that applies sound to the video.

An AI-generated video from Veo 3: "A male stand-up comic on stage in a night club telling a hilarious joke about AI and crypto with a silly punchline." An AI language model built into Veo 3 wrote the joke.

An AI-generated video from Veo 3: "A male stand-up comic on stage in a night club telling a hilarious joke about AI and crypto with a silly punchline." An AI language model built into Veo 3 wrote the joke.

In an attempt to prevent misuse, DeepMind says it's using its proprietary watermarking technology, SynthID, to embed invisible markers into frames Veo 3 generates. These watermarks persist even when videos are compressed or edited, helping people potentially identify AI-generated content. As we'll discuss more later, though, this may not be enough to prevent deception.

Google also censors certain prompts and outputs that breach the company's content agreement. During testing, we encountered "generation failure" messages for videos that involve romantic and sexual material, some types of violence, mentions of certain trademarked or copyrighted media properties, some company names, certain celebrities, and some historical events.

Putting Veo 3 to the test

Perhaps the biggest change with Veo 3 is integrated audio generation, although Meta previewed a similar audio-generation capability with "Movie Gen" last October, and AI researchers have experimented with using AI to add soundtracks to silent videos for some time. Google DeepMind itself showed off an AI soundtrack-generating model in June 2024.

An AI-generated video from Veo 3: "A middle-aged balding man rapping indie core about Atari, IBM, TRS-80, Commodore, VIC-20, Atari 800, NES, VCS, Tandy 100, Coleco, Timex-Sinclair, Texas Instruments"

An AI-generated video from Veo 3: "A middle-aged balding man rapping indie core about Atari, IBM, TRS-80, Commodore, VIC-20, Atari 800, NES, VCS, Tandy 100, Coleco, Timex-Sinclair, Texas Instruments"

Veo 3 can generate everything from traffic sounds to music and character dialogue, though our early testing reveals occasional glitches. Spaghetti makes crunching sounds when eaten (as we covered last week, with a nod to the famous Will Smith AI spaghetti video), and in scenes with multiple people, dialogue sometimes comes from the wrong character's mouth. But overall, Veo 3 feels like a step change in video synthesis quality and coherency over models from OpenAI, Runway, Minimax, Pika, Meta, Kling, and Hunyuanvideo.

The videos also tend to show garbled subtitles that almost match the spoken words, which is an artifact of subtitles on videos present in the training data. The AI model is imitating what it has "seen" before.

An AI-generated video from Veo 3: "A beer commercial for 'CATNIP' beer featuring a real a cat in a pickup truck driving down a dusty dirt road in a trucker hat drinking a can of beer while country music plays in the background, a man sings a jingle 'Catnip beeeeeeeeeeeeeeeeer' holding the note for 6 seconds"

An AI-generated video from Veo 3: "A beer commercial for 'CATNIP' beer featuring a real a cat in a pickup truck driving down a dusty dirt road in a trucker hat drinking a can of beer while country music plays in the background, a man sings a jingle 'Catnip beeeeeeeeeeeeeeeeer' holding the note for 6 seconds"

We generated each of the eight-second-long 720p videos seen below using Google's Flow platform. Each video generation took around three to five minutes to complete, and we paid for them ourselves. It's important to note that better results come from cherry-picking—running the same prompt multiple times until you find a good result. Due to cost and in the spirit of testing, we only ran every prompt once, unless noted.

New audio prompts

Let's dive right into the deep end with audio generation to get a grip on what this technology can do. We've previously shown you a man singing about spaghetti and a rapping shark in our last Veo 3 piece, but here's some more complex dialogue.

Since 2022, we've been using the prompt "a muscular barbarian with weapons beside a CRT television set, cinematic, 8K, studio lighting" to test AI image generators like Midjourney. It's time to bring that barbarian to life.

A muscular barbarian man holding an axe, standing next to a CRT television set. He looks at the TV, then to the camera and literally says, "You've been looking for this for years: a muscular barbarian with weapons beside a CRT television set, cinematic, 8K, studio lighting. Got that, Benj?"

The video above represents significant technical progress in AI media synthesis over the course of only three years. We've gone from a blurry colorful still-image barbarian to a photorealistic guy that talks to us in 720p high definition with audio. Most notably, there's no reason to believe technical capability in AI generation will slow down from here.

Horror film: A scared woman in a Victorian outfit running through a forest, dolly shot, being chased by a man in a peanut costume screaming, "Wait! You forgot your wallet!"

Trailer for The Haunted Basketball Train: a Tim Burton film where 1990s basketball star is stuck at the end of a haunted passenger train with basketball court cars, and the only way to survive is to make it to the engine by beating different ghosts at basketball in every car

ASMR video of a muscular barbarian man whispering slowly into a microphone, "You love CRTs, don't you? That's OK. It's OK to love CRT televisions and barbarians."

1980s PBS show about a man with a beard talking about how his Apple II computer can "connect to the world through a series of tubes"

A 1980s fitness video with models in leotards wearing werewolf masks

A female therapist looking at the camera, zoom call. She says, "Oh my lord, look at that Atari 800 you have behind you! I can't believe how nice it is!"

With this technology, one can easily imagine a virtual world of AI personalities designed to flatter people. This is a fairly innocent example about a vintage computer, but you can extrapolate, making the fake person talk about any topic at all. There are limits due to Google's filters, but from what we've seen in the past, a future uncensored version of a similarly capable AI video generator is very likely.

Video call screenshot capture of a Zoom chat. A psychologist in a dark, cozy therapist's office. The therapist says in a friendly voice, "Hi Tom, thanks for calling. Tell me about how you're feeling today. Is the depression still getting to you? Let's work on that."

1960s NASA footage of the first man stepping onto the surface of the Moon, who squishes into a pile of mud and yells in a hillbilly voice, "What in tarnation??"

A local TV news interview of a muscular barbarian talking about why he's always carrying a CRT TV set around with him

Speaking of fake news interviews, Veo 3 can generate plenty of talking anchor-persons, although sometimes on-screen text is garbled if you don't specify exactly what it should say. It's in cases like this where it seems Veo 3 might be most potent at casual media deception.

Footage from a news report about Russia invading the United States

Attempts at music

Veo 3's AI audio generator can create music in various genres, although in practice, the results are typically simplistic. Still, it's a new capability for AI video generators. Here are a few examples in various musical genres.

A PBS show of a crazy barbarian with a blonde afro painting pictures of Trees, singing "HAPPY BIG TREES" to some music while he paints

A 1950s cowboy rides up to the camera and sings in country music, "I love mah biiig ooold donkeee"

A 1980s hair metal band drives up to the camera and sings in rock music, "Help me with my huge huge huge hair!"

Mister Rogers' Neighborhood PBS kids show intro done with psychedelic acid rock and colored lights

1950s musical jazz group with a scat singer singing about pickles amid gibberish

Some classic prompts from prior tests

The prompts below come from our previous video tests of Gen-3, Video-01, and the open source Hunyuanvideo, so you can flip back to those articles and compare the results if you want to. Overall, Veo 3 appears to have far greater temporal coherency (having a consistent subject or theme over time) than the earlier video synthesis models we've tested. But of course, it's not perfect.

A highly intelligent person reading 'Ars Technica' on their computer when the screen explodes

The moonshark jumping out of a computer screen and attacking a person

A herd of one million cats running on a hillside, aerial view

Video game footage of a dynamic 1990s third-person 3D platform game starring an anthropomorphic shark boy

Aerial shot of a small American town getting deluged with liquid cheese after a massive cheese rainstorm where liquid cheese rained down and dripped all over the buildings

Wide-angle shot, starting with the Sasquatch at the center of the stage giving a TED talk about mushrooms, then slowly zooming in to capture its expressive face and gestures, before panning to the attentive audience

A trip-hop rap song about Ars Technica being sung by a guy in a large rubber shark costume on a stage with a full moon in the background

Some notable failures

Google's Veo 3 isn't perfect at synthesizing every scenario we can throw at it due to limitations of training data. As we noted in our previous coverage, AI video generators remain fundamentally imitative, making predictions based on statistical patterns rather than a true understanding of physics or how the world works.

For example, if you see mouths moving during speech, or clothes wrinkling in a certain way when touched, it means the neural network doing the video generation has "seen" enough similar examples of that scenario in the training data to render a convincing take on it and apply it to similar situations.

However, when a novel situation (or combination of themes) isn't well-represented in the training data, you'll see "impossible" or illogical things happen, such as weird body parts, magically appearing clothing, or an object that "shatters" but remains in the scene afterward, as you'll see below.

We mentioned audio and video glitches in the introduction. In particular, scenes with multiple people sometimes confuse which character is speaking, such as this argument between tech fans.

A 2000s TV debate between fans of the PowerPC and Intel Pentium chips

Bombastic 1980s infomercial for the "Ars Technica" online service. With cheesy background music and user testimonials

1980s Rambo fighting Soviets on the Moon

Sometimes requests don't make coherent sense. In this case, "Rambo" is correctly on the Moon firing a gun, but he's not wearing a spacesuit. He's a lot tougher than we thought.

An animated infographic showing how many floppy disks it would take to hold an installation of Windows 11

Large amounts of text also present a weak point, but if a short text quotation is explicitly specified in the prompt, Veo 3 usually gets it right.

A young woman doing a complex floor gymnastics routine at the Olympics, featuring running and flips

Despite Veo 3's advances in temporal coherency and audio generation, it still suffers from the same "jabberwockies" we saw in OpenAI's viral Sora gymnast video—those non-plausible video hallucinations like impossible morphing body parts.

A silly group of men and women cartwheeling across the road, singing "CHEEEESE" and holding the note for 8 seconds before falling over.

A YouTube-style try-on video of a person trying on various corncob costumes. They shout "Corncob haul!!"

A man made of glass runs into a brick wall and shatters, screaming

A man in a spacesuit holding up 5 fingers and counting down to zero, then blasting off into space with rocket boots

Counting down with fingers is difficult for Veo 3, likely because it's not well-represented in the training data. Instead, hands are likely usually shown in a few positions like a fist, a five-finger open palm, a two-finger peace sign, and the number one.

As new architectures emerge and future models train on vastly larger datasets with exponentially more compute, these systems will likely forge deeper statistical connections between the concepts they observe in videos, dramatically improving quality and also the ability to generalize more with novel prompts.

The “cultural singularity” is coming—what more is left to say?

By now, some of you might be worried that we're in trouble as a society due to potential deception from this kind of technology. And there's a good reason to worry: The American pop culture diet currently relies heavily on clips shared by strangers through social media such as TikTok, and now all of that can easily be faked, whole-cloth. Automated generations of fake people can now argue for ideological positions in a way that could manipulate the masses.

AI-generated video by Veo 3: "A man on the street interview about someone who fears they live in a time where nothing can be believed"

AI-generated video by Veo 3: "A man on the street interview about someone who fears they live in a time where nothing can be believed"

Such videos could be (and were) manipulated before through various means prior to Veo 3, but now the barrier to entry has collapsed from requiring specialized skills, expensive software, and hours of painstaking work to simply typing a prompt and waiting three minutes. What once required a team of VFX artists or at least someone proficient in After Effects can now be done by anyone with a credit card and an Internet connection.

But let's take a moment to catch our breath. At Ars Technica, we've been warning about the deceptive potential of realistic AI-generated media since at least 2019. In 2022, we talked about AI image generator Stable Diffusion and the ability to train people into custom AI image models. We discussed Sora "collapsing media reality" and talked about persistent media skepticism during the "deep doubt era."

AI-generated video with Veo 3: "A man on the street ranting about the 'cultural singularity' and the 'cultural apocalypse' due to AI"

AI-generated video with Veo 3: "A man on the street ranting about the 'cultural singularity' and the 'cultural apocalypse' due to AI"

I also wrote in detail about the future ability for people to pollute the historical record with AI-generated noise. In that piece, I used the term "cultural singularity" to denote a time when truth and fiction in media become indistinguishable, not only because of the deceptive nature of AI-generated content but also due to the massive quantities of AI-generated and AI-augmented media we'll likely soon be inundated with.

However, in an article I wrote last year about cloning my dad's handwriting using AI, I came to the conclusion that my previous fears about the cultural singularity may be overblown. Media has always been vulnerable to forgery since ancient times; trust in any remote communication ultimately depends on trusting its source.

AI-generated video with Veo 3: "A news set. There is an 'Ars Technica News' logo behind a man. The man has a beard and a suit and is doing a sit-down interview. He says "This is the age of post-history: a new epoch of civilization where the historical record is so full of fabrication that it becomes effectively meaningless."

AI-generated video with Veo 3: "A news set. There is an 'Ars Technica News' logo behind a man. The man has a beard and a suit and is doing a sit-down interview. He says "This is the age of post-history: a new epoch of civilization where the historical record is so full of fabrication that it becomes effectively meaningless."

The Romans had laws against forgery in 80 BC, and people have been doctoring photos since the medium's invention. What has changed isn't the possibility of deception but its accessibility and scale.

With Veo 3's ability to generate convincing video with synchronized dialogue and sound effects, we're not witnessing the birth of media deception—we're seeing its mass democratization. What once cost millions of dollars in Hollywood special effects can now be created for pocket change.

An AI-generated video created with Google Veo-3: "A candid interview of a woman who doesn't believe anything she sees online unless it's on Ars Technica."

An AI-generated video created with Google Veo-3: "A candid interview of a woman who doesn't believe anything she sees online unless it's on Ars Technica."

As these tools become more powerful and affordable, skepticism in media will grow. But the question isn't whether we can trust what we see and hear. It's whether we can trust who's showing it to us. In an era where anyone can generate a realistic video of anything for $1.50, the credibility of the source becomes our primary anchor to truth. The medium was never the message—the messenger always was.