Meta launches Muse Image, previews Muse Video
By Steven Van ·
Muse Image plans and edits like an agent and is live in Meta AI, Instagram Stories and WhatsApp, while Muse Video is coming soon.
Meta Superintelligence Labs has released Muse Image, which Meta calls its most advanced image generation model yet, and gave a first look at Muse Video, a video model built on the same foundation. The announcement lands two months after Muse Spark, MSL's first model, and it signals how quickly Meta is turning its superintelligence lab into a consumer-facing model factory: Muse Image is not a research preview, it's live today inside the Meta AI app, on meta.ai, in Instagram Stories in the US, and in WhatsApp in limited countries, with Facebook support coming soon.
An image model that thinks before it draws
The interesting part of Muse Image is not the pixels, it's the process. Instead of mapping a prompt directly to an image, the model behaves like an agent. Paired with Muse Spark, it plans the layout of a scene before rendering it, invokes tools along the way, and refines its own output over multiple passes, a behavior Meta says emerged during training and improves with test-time compute.
The tool use is concrete rather than hand-wavy. Muse Image can run code to produce plots, QR codes, and animated GIFs inside a generation, and it can search the web to ground an image in facts. Ask for an infographic about a real subject and it can look the subject up first, then render the result with clean, legible, styled text, historically one of the most reliable ways to make an image model embarrass itself.
Editing and multi-reference composition
Beyond fresh generations, Muse Image is built for precise editing that holds up across multiple refinement turns, so you can keep adjusting one image conversationally without it degrading. It also composes from several references at once, blending people, objects, clothing, and styles from different input images into a single coherent output. In Meta's apps this shows up as presets that suggest ideas and the ability to sketch edits directly on an image.
The benchmarks back the pitch: as of early July, Muse Image holds the No. 2 spot on Arena's human-preference leaderboards for text-to-image, single-image editing, and multi-image editing.
Watermarking, and a privacy flashpoint
Every Muse Image generation carries Content Seal, Meta's invisible watermark, so AI-generated images can be verified after the fact. The launch has not been friction-free, though. TechCrunch reports that users are already pushing back over the model drawing on their photos and Instagram activity for what Meta calls social context, personalization that some see as the feature and others see as the problem. If you create with it, it's worth knowing both halves of that trade.
Muse Video is next
Muse Video, built on the same pretraining foundation, is Meta's answer to the text-to-video race. Meta says it delivers exceptional visual fidelity with native audio support, meaning sound is generated with the video rather than bolted on, and it currently ranks No. 3 on Arena for text-to-video. It isn't public yet: Meta says it's coming soon to creators and to Meta AI. For creators, the combined pitch is clear enough, one model family handling on-brand images, edits, and eventually video, inside the apps where the audience already is.
Sources: Meta AI, Meta Newsroom, TechCrunch.