Metaverse and A.I.

OmniHuman: ByteDance’s New AI Creates Realistic Videos From a Single Photo

ByteDance researchers have developed an AI system that transforms single photographs into realistic videos of people speaking, singing and moving naturally — a breakthrough that could reshape digital entertainment and communications. The new system, called OmniHuman, generates full-body videos that show people gesturing and moving in ways that match their speech, surpassing previous AI models that could only animate faces or upper bodies.

“End-to-end human animation has undergone notable advancements in recent years,” the ByteDance researchers wrote in a paper published on arXiv. “However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications,”  The team trained OmniHuman on more than 18,700 hours of human video data using a novel approach that combines multiple types of inputs — text, audio and body movements. This “omni-conditions” training strategy allows the AI to learn from much larger and more diverse datasets than previous methods.

“Our key insight is that incorporating multiple conditioning signals, such as text, audio and pose, during training can significantly reduce data wastage,” the research team explained. The technology marks a significant advance in AI-generated media, demonstrating capabilities that range from creating videos of people delivering speeches to depicting subjects playing musical instruments. In testing, OmniHuman outperformed existing systems across multiple quality benchmarks.

The development emerges amid intensifying competition in AI video generation, with companies like GoogleMeta and Microsoft pursuing similar technologies. ByteDance’s breakthrough could give its TikTok parent company an advantage in this rapidly evolving field. Industry experts say such technology could transform entertainment production, educational content creation and digital communications. However, it also raises concerns about potential misuse in creating synthetic media for deceptive purposes. The researchers will present their findings at an upcoming computer vision conference, although they have not yet specified when or which one.

Terron Gold

Recent Posts

Elon Musk’s xAI Sues Minnesota Over First U.S. AI Nudification Law

Elon Musk's artificial intelligence company xAI has filed a federal lawsuit challenging Minnesota's landmark law banning AI-powered "nudification" technology, arguing…

6 days ago

Bitcoin Nears $65,000 as Treasury Yields Outperform Carry Trade in Rare Market Signal

Bitcoin climbed toward $65,000 as an unusual shift in traditional financial markets created one of the rarest conditions…

7 days ago

Hyperscale Data Sells 100 Bitcoin to Fund Michigan AI Campus as GPUS Stock Surges

Hyperscale Data has sold 100 Bitcoin to accelerate development of its planned artificial intelligence campus in Michigan, sending shares…

7 days ago

PIPEDOG Explodes 140x on Robinhood Chain as Memecoin Frenzy Intensifies

A newly launched memecoin called PIPEDOG ($PIPEDOG) became the latest breakout token on Robinhood Chain, surging more than 140x within…

1 week ago

BNY Brings $8.6 Trillion Fund Business On-Chain in Major Wall Street Blockchain Expansion

BNY, the world's largest custodian bank, is bringing one of its core financial businesses onto…

1 week ago

Senators Strengthen Crypto Ethics Rules in CLARITY Act After Trump Negotiations

A bipartisan group of U.S. senators has reportedly reached a new compromise on the ethics provisions of…

1 week ago