What AI Image and Video Generation Can Actually Do
A capability-honest guide to text-to-image, text-to-video, and everything around them as of September 2026 — what still breaks, how the three pricing models actually charge you, and the legal checklist to run before you generate anything commercial.
Every few months, a new model launch makes AI image and video generation look like it has solved everything. It hasn’t, and the gap between “the demo reel” and “what happens the fortieth time you try to get a specific shot” is where most builders and small businesses lose money and time. This guide is not a product recommendation — for “which specific tool should I use,” see Midjourney vs DALL-E vs Stable Diffusion for images and Sora vs Runway vs Veo for video, both current as of September 2026 and cited directly here rather than re-derived. What follows instead is the layer underneath those choices: what current models genuinely can and can’t do, how the pricing actually works once you look past the sticker price, and the checklist — legal, contractual, and practical — worth running before you generate anything you intend to publish or sell.
The capability landscape, honestly
Images
Text-to-image is the mature case. Type a description, get a picture. Every major model handles this well enough that the interesting differences are now in style, instruction-following, and cost rather than whether it works at all.
Image-to-image and editing — taking an existing image and modifying it via a new prompt rather than starting from noise — has moved from a niche feature to a default workflow. Conversational editing (say what to change, get an updated image in the same thread, without re-writing the whole prompt) is now how OpenAI’s GPT Image 2 works inside ChatGPT, and Google’s Nano Banana Pro is built around the same interaction model with a specific strength in accurate multilingual text rendering.
Inpainting and outpainting — regenerating a masked region of an image, or extending a canvas beyond its original borders — are standard features across every major tool now, used for object removal, background extension, and compositing generated elements into existing photography.
Reference-image and style-guided generation covers two related but distinct things. ControlNet-style conditioning constrains a generation to a specific pose, depth map, or edge layout — strongest in the Stable Diffusion ecosystem, where ControlNet is a first-class tool, and comparatively weak or absent in Midjourney and GPT Image. Character and style consistency across a series is the harder problem: reference-image pinning holds up reasonably well within a similar pose and degrades once you ask for the same character in a substantially different pose or scene. Fine-tuning your own LoRA on a specific character remains the more reliable route to a genuinely consistent series — a route only open on Stable Diffusion; see Midjourney vs DALL-E vs Stable Diffusion for how that fits into the ecosystem.
Upscaling is a solved, unglamorous feature at this point — every major tool offers it, and it’s rarely the differentiator between products.
What’s genuinely still broken, honestly: complex hand poses — several fingers interacting, hands holding small objects — still produce occasional artifacts, though simple hand poses are now handled reliably, a real change from 2023’s near-universal hand problem. Text rendering inside images was a much bigger weakness a year ago; Nano Banana Pro and GPT Image 2 have made legible, accurately-spelled in-image text a realistic expectation rather than a lucky outcome, though not a guaranteed one. Physical plausibility — lighting that doesn’t match its source, shadows falling the wrong way, objects interpenetrating — has improved the least, because it requires reasoning about physics the model was never explicitly taught, not just more pattern-matching.
Video
Text-to-video and image-to-video (animating a still into motion) are both standard across Runway, Veo, and the tools built on top of them. Image-to-video is generally the more reliable of the two for a specific look, since a real starting frame constrains the model more than a text description alone.
Video extension and continuation is how every current model gets past its native clip-length ceiling: the platform takes the final frame of a clip, treats it as a fresh starting image, and generates a new segment from there. Native single-pass generation runs roughly 5–30 seconds across the major models (Veo 3.1 generates 8 seconds natively; Runway’s Gen-4.5 API generates 2–10 seconds per call), and extension chains push well past that — Veo’s mechanism can reach roughly 148 seconds through repeated 7-second extensions, Kling’s chaining reaches several minutes. The catch: each extension is conditioned only on the last frame, not the video’s full history, so drift accumulates — a character’s face, outfit, or the scene’s lighting can subtly shift over a chain in a way it wouldn’t within one native generation.
Native audio and lip-sync generation is a genuine, recent capability split, not a universal feature. Veo 3.1 generates dialogue, sound effects, and ambient audio in the same pass — no separate audio tool required — but Google’s own Veo documentation states plainly that “creating videos with natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development”: real, not fully reliable, especially for lip-synced dialogue. Runway’s Gen-4.5 does not clearly generate native audio inside the base model itself; sources disagree on whether that’s since changed, and Runway’s own docs don’t settle it — see Sora vs Runway vs Veo for the detail. Check current documentation per project rather than assuming either way.
Camera-control and motion-direction features are where the tools differentiate the most from each other right now. Runway’s Director Mode gives parametric camera moves — dolly, orbit, crane, rack focus — without a physical rig, and its Motion Brush lets you paint motion vectors onto specific regions of a frame to control exactly what moves. Nothing in Veo’s public feature set matches this level of granular camera control as of this writing.
What’s genuinely still broken, honestly: character and scene consistency across multiple shots or a long video is the single hardest unsolved problem in this category. Models generate each shot with limited or no memory of the previous one, so even a clear text description of “the same person” can drift in facial features, hairstyle, or proportions across cuts — degradation that gets noticeably worse past roughly 30 seconds of cumulative generated footage on most current models, and worse still the more a new shot’s pose or angle diverges from the reference. Some 2026-generation models (Kling 3.0’s multi-shot storyboard mode is the most cited example) are starting to address this at the model level rather than leaving it to prompt engineering, but it is not a solved problem the way hands-in-still-images increasingly is — this is one of the limitations moving slowly, not one that’s been substantially fixed in the last year. Physical plausibility failures carry over from images and compound over time in video: objects passing through each other, inconsistent shadows, motion that doesn’t obey momentum.
Pricing models explained
Nearly every image and video generation product on the market bills through one of three models. None of them price by “the image” or “the video” in a simple, comparable way — understanding the mechanics matters more than memorizing any one vendor’s sticker price, which is why the current numbers live in the vendor-specific comparisons rather than here.
| Model | How you pay | What actually drives cost per output | Typical fallback when you run out |
|---|---|---|---|
| Credit-based subscription | Monthly fee buys a pool of credits; each generation consumes a variable number depending on settings | Resolution, clip duration (video), model tier (turbo/fast vs. pro/max), export format (ProRes, HDR, and near-4K formats carry real credit surcharges on Runway, for example) | Buy more credits, or wait for the monthly refill |
| Flat-tier subscription with a slow fallback | Flat monthly fee for a tier; “fast” generation is capped, then a slower “relaxed” mode kicks in | Which tier you’re on (higher tiers get more fast-generation hours, not unlimited fast generation); resolution and job complexity still affect how many fast-hours a given output consumes | Unlimited but queued/slower generation — Midjourney is the clearest example of this model |
| Pure metered API pricing | Pay per unit consumed: per second of video, per output token (images), or per API call | Resolution and quality setting, output duration, whether you’re hitting a “fast/standard” or “pro/ultra” model variant, batch vs. real-time processing (batch APIs commonly run at roughly half the real-time rate) | None — the meter just keeps running; this is the model developer-facing platforms use almost universally |
A few vendor-specific numbers, all cited from this wiki’s dedicated comparisons rather than re-verified here: Midjourney’s Basic plan ($10/month) buys roughly 200 fast images before falling back to unlimited relaxed generation; OpenAI’s GPT Image 2 API bills per output token, working out to roughly $0.02–$0.20 per image depending on resolution and quality; Runway’s Gen-4.5 API runs $0.12/second of video, with HDR and near-4K formats adding substantial per-second surcharges on top; Google’s Veo 3.1 API ranges $0.05–$0.60/second depending on tier and resolution. See the two comparison pages linked above for the full current tables and how each vendor’s consumer subscription tiers translate into those API rates.
What none of these sticker prices tell you: your real yield. Prompting is a generate-many-discard-most process — professional workflows routinely generate five, ten, or more variations to get one usable output, and that ratio is invisible in a per-image or per-second price. The number that actually matters for budgeting is cost-per-usable output, not cost-per-generation: divide your real spend for a project by the number of outputs you actually kept and used, not the number you generated. A $0.05 image sounds cheap until a project’s real yield is one keeper in eight attempts, which makes the effective cost $0.40 — worth knowing before committing to a volume estimate, and worth tracking retroactively on your first few real projects with any given tool so later estimates aren’t just guesses. Retries and near-miss regenerations are the single biggest hidden cost in every pricing model above; none of the three structurally discounts for them.
What to check before you generate
Consent and likeness rights
Generating a real, identifiable person’s likeness without consent is a distinct source of legal exposure, separate from any AI-specific regulation and separate again from labeling obligations. In the US, this sits under state-level right-of-publicity law, which varies significantly — Tennessee’s ELVIS Act (in effect since July 2024) was among the first to explicitly extend right-of-publicity protection to AI voice replication, and Washington state added “forged digital likenesses” to its personality-rights framework in a law effective June 2026. Federally, the bipartisan NO FAKES Act — a proposed right to authorize use of one’s voice or likeness in AI “digital replicas,” preempting some inconsistent state laws — advanced unanimously out of the Senate Judiciary Committee in June 2026 but had not been enacted as of this writing. In the EU, personality and image rights sit at member-state level, layered on top of, not replaced by, the AI Act’s own disclosure duty (below). The two questions are genuinely separate: labeling a deepfake as AI-generated, as EU law requires, is not permission to have generated a real person’s likeness in the first place, and doing so without consent can create right-of-publicity or personality-rights exposure regardless of labeling. This isn’t theoretical — see Deepfake Fraud Reaches $3.7 Billion in Documented Losses for how far unauthorized likeness generation has already been weaponized.
Copyright and training-data provenance
Whether training an image or video model on copyrighted material is infringement remains genuinely unsettled. In Getty Images v. Stability AI, the UK High Court substantially dismissed Getty’s claims in November 2025 — Getty had dropped its primary training-data infringement claim for jurisdictional reasons, and the court found no secondary copyright infringement, though it did find limited, “historic and extremely limited” trademark infringement from Getty watermarks appearing in outputs of earlier Stable Diffusion versions. Getty was granted permission to appeal the secondary-infringement finding in December 2025; that appeal is pending. In parallel, a US federal case over the same underlying dispute (N.D. Cal.) survived Stability’s motion to dismiss on the core copyright, trademark, and unfair-competition claims — so the training-data question is being actively litigated on two continents with different, not-yet-final outcomes on each.
Separately, and more settled: the US Copyright Office’s January 2025 report on AI copyrightability concluded that output generated with no meaningful human creative control is not eligible for copyright protection, and specifically that text prompts alone — however detailed — don’t provide enough control to establish human authorship, since the same prompt can produce widely different results. A work combining human-authored elements with AI-generated material, or where a human creatively arranges or modifies AI output, can still be copyrightable in the human-authored parts. This was reinforced by the Supreme Court’s March 2026 denial of certiorari in Thaler v. Perlmutter, leaving the DC Circuit’s human-authorship requirement as settled US baseline. The practical upshot: content you generate with minimal creative input may not be protectable by you against a competitor who copies it outright — a separate risk from whether your output itself infringes someone else’s copyright.
Output ownership and commercial usage rights
These differ substantially by vendor and, within a vendor, by tier — verify current terms directly before assuming, because the differences are real and checkable, not boilerplate.
| Vendor | Who owns the output | Commercial use gating | Notes |
|---|---|---|---|
| Midjourney | You, per its Terms of Service | All paid tiers get commercial rights — but a business earning over $1M/year gross revenue must specifically be on the Pro plan to keep them | No free tier exists at all; see Midjourney vs DALL-E vs Stable Diffusion |
| OpenAI (GPT Image 2 / ChatGPT) | You — OpenAI assigns its rights in output to the customer | Commercial use permitted on both free and paid tiers | Standard terms disclaim any warranty that output is non-infringing |
| Google (Gemini API) | You — Google’s Additional Terms of Service state Google does not claim ownership of generated output | Broader commercial rights apply on paid API access than on unpaid/consumer use | Google also reserves the right to generate similar content for other users |
| Stability AI (Stable Diffusion) | You, under the Community License | Free for commercial use under roughly $1M/year revenue; an Enterprise License is required above that threshold | The only one of these you can also run entirely offline, which sidesteps the licensing terms of a hosted API entirely |
| Runway | You — Runway does not claim ownership of your inputs or outputs | Paid subscribers get commercial rights; free-tier use is personal/non-commercial only | Runway’s terms reportedly retain a license to use your content for model training except on Enterprise plans — confirm current wording before assuming otherwise |
Indemnification
Indemnification — a vendor agreeing to cover your legal costs and damages if a generated output triggers a third-party IP claim — is genuine, checkable, and unevenly distributed. Adobe offers the most explicit commitment: it will defend a claim that a Firefly output directly infringes copyright, trademark, or publicity/privacy rights, and pay resulting damages or settlements — but only for Creative Cloud for Teams or Enterprise customers on specific plans (Pro Edition or Edition 4) that include Firefly Output Indemnification; it doesn’t extend to free or standard individual tiers, and excludes claims arising from your own modifications to the output. OpenAI extends indemnification to Enterprise customers as part of its enterprise agreement terms, not to consumer or self-serve API terms. Google made a similar generative-AI indemnification pledge for enterprise Google Cloud customers back in 2023, but its Gemini API Additional Terms of Service — what actually governs most developers calling the API directly — contain no indemnification clause at all, worth flagging precisely because it’s easy to assume a company-wide pledge covers a product it doesn’t. Midjourney and Stability AI’s standard terms disclaim IP warranties and offer no indemnification at any tier; Runway’s available terms don’t clearly document a standard-tier offer either. The pattern: where indemnification exists, it’s enterprise-tier or product-scoped, not a blanket benefit of any paid subscription — confirm the specific plan you’re on, not the vendor’s general marketing claim.
Content policy limits
Every major vendor restricts some category of generation, and the restrictions differ enough to matter for a specific use case. GPT Image 2 inherits ChatGPT’s broader content policy and is noticeably more conservative than either Midjourney or an uncensored local Stable Diffusion checkpoint on real people’s likenesses, some copyrighted characters, and other edge-case requests. Midjourney’s Community Guidelines don’t flatly ban generating real people but prohibit using such imagery to harass, abuse, defame, or otherwise harm someone, alongside separate bans on election-related content intended to influence outcomes and any sexualized content involving minors. If a project’s viability depends on a specific category of output — real public figures, particular content styles, specific edge cases — check the current policy of the specific tool and tier directly; these change without much notice and differ meaningfully between products that otherwise look interchangeable.
Provenance and watermarking
Marking your outputs as AI-generated matters beyond any regulatory requirement to do so: downstream platforms increasingly require disclosure as a matter of their own policy, independent of what law requires in your jurisdiction, and unmarked synthetic content erodes trust in ways that outlast any single post or campaign. This wiki’s AI Watermarking glossary entry covers how the underlying technical mechanisms actually work — statistical signatures embedded in the generation process, and how they survive or fail to survive common transformations; it’s not repeated here. What’s worth adding in a builder-facing context: watermarking is one layer of a provenance strategy, not the whole thing, and it works best paired with metadata standards like C2PA and your own generation logs, particularly if you’ll ever need to demonstrate that a specific piece of content was or wasn’t generated by your system.
How this connects to the EU AI Act
Article 50 of the EU AI Act (Regulation (EU) 2024/1689) splits the transparency obligation for synthetic image and video content into two distinct duties on two distinct parties, and conflating them is a common mistake. Article 50(2) puts a marking duty on providers: anyone whose system generates synthetic audio, image, video, or text content must ensure the output is marked in a machine-readable format that’s detectable as artificially generated, with narrow exceptions for systems performing an assistive editing function that doesn’t substantially alter input data. Article 50(4) puts a separate, human-facing disclosure duty on deployers: whoever uses an AI system to generate or manipulate content that meets the Act’s definition of a “deep fake” — defined in Article 3(60) as AI-generated or manipulated image, audio, or video content that resembles a real person, object, place, entity, or event and would falsely appear authentic — must disclose that the content was artificially generated or manipulated to any natural person exposed to it, at first exposure at the latest, in a way that’s clear, distinguishable, and doesn’t depend on the viewer already having Article 50(2)’s machine-readable marking or any special tooling to detect it.
Both obligations have applied, enforceably, since 2 August 2026, and — unlike the Act’s Annex III high-risk system obligations, which the Digital Omnibus deferred to December 2027 — Article 50 was not part of that deferral. One clarification worth having precise: the deepfake disclosure duty applies regardless of intent to deceive, so content that resembles a real person triggers it even where nobody involved meant to mislead anyone. There’s also a real carve-out for creative work: where the content is evidently part of an artistic, creative, satirical, or fictional work, the obligation narrows to disclosing that generated or manipulated content exists somewhere in the work, in a way that doesn’t get in the way of experiencing it — not a full first-exposure disclosure of every synthetic element.
This is deliberately the short version. For the fuller mechanics — how the marking and disclosure duties interact in practice, what “clear and distinguishable” has been interpreted to require, how this obligation sits alongside the rest of the Act’s transparency regime, and the practical steps a team actually generating image or video content at scale needs to take — see Synthetic Media and the EU AI Act’s Transparency Rules , and Deepfake for the term’s precise definition and how it’s used across the regulation.
Sources
- Google DeepMind, official Veo model page, including the quoted limitation on spoken-audio consistency: https://deepmind.google/models/veo/
- Runway Developer Docs, “API Pricing & Costs” (credit rates, professional/HDR format surcharges): https://docs.dev.runwayml.com/guides/pricing/
- Google AI for Developers, Gemini API Additional Terms of Service (output ownership, absence of an indemnification clause, paid vs. unpaid data-use terms): https://ai.google.dev/gemini-api/terms
- U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (29 January 2025): https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf
- Mayer Brown, “Supreme Court Denies Cert in AI Authorship Case” (Thaler v. Perlmutter, cert denied 2 March 2026): https://www.mayerbrown.com/en/insights/publications/2026/03/supreme-court-denies-review-in-ai-authorship-case
- Loeb & Loeb, “Getty Images (US), Inc. v. Stability AI, Ltd.” (case status summary, UK and US proceedings, April 2026): https://www.loeb.com/en/insights/publications/2026/04/getty-images-us-inc-v-stability-ai-ltd
- Bloomberg Law, “Getty’s AI Copyright Suit Survives Stability’s Bid for Dismissal” (US case, N.D. Cal.): https://news.bloomberglaw.com/ip-law/gettys-ai-copyright-suit-survives-stabilitys-bid-for-dismissal
- Getty Images Newsroom, official statement on the UK High Court ruling: https://newsroom.gettyimages.com/en/getty-images/getty-images-issues-statement-on-ruling-in-stability-ai-uk-litigation
- Computerworld, “Adobe offers copyright indemnification for Firefly AI-based image app users” (indemnification scope by Creative Cloud plan tier): https://www.computerworld.com/article/1628682/adobe-offers-copyright-indemnification-for-firefly-ai-based-image-app-users.html
- Adobe, Firefly business overview: https://business.adobe.com/products/firefly-business/firefly-ai-approach.html
- Runtime, “AI vendors promised indemnification against copyright lawsuits. The details are messy.” (vendor-by-vendor indemnification conditions and exclusions): https://www.runtime.news/ai-vendors-promised-indemnification-against-copyright-lawsuits-the-details-are-messy/
- terms.law, Midjourney commercial-use policy summary (the $1M revenue / Pro-plan requirement for retaining commercial rights): https://terms.law/ai-output-rights/midjourney/
- Stability AI, Community License terms ($1M annual revenue threshold for free commercial use): https://stability.ai/license
- artificialintelligenceact.eu, Article 50 full text and practical guide, including the deepfake definition (Article 3(60)) and the artistic-work exception: https://artificialintelligenceact.eu/article/50/ , https://artificialintelligenceact.eu/transparency-rules-article-50/
- Regulation (EU) 2024/1689 (AI Act), Article 50, EUR-Lex: https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- European Commission, FAQs on transparency obligations under Article 50: https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act
- ByteBack Law, “A Federal Shift in AI and the Right of Publicity? NO FAKES Act Advances in Congress” (June 2026 committee vote): https://www.bytebacklaw.com/2026/08/a-federal-shift-in-ai-and-the-right-of-publicity-no-fakes-act-advances-in-congress/
- Holland & Knight, “Senate Committee Advances Legislation to Protect Name, Image, Likeness” (NO FAKES Act, Washington state ELVIS-style law): https://www.hklaw.com/en/insights/publications/2026/06/senate-judiciary-committee-advances-legislation-to-protect-name
Further reading
- Midjourney vs DALL-E vs Stable Diffusion : current pricing, positioning, and which image tool fits which use case.
- Sora vs Runway vs Veo : current pricing, camera-control and native-audio detail, and which video tool fits which use case.
- Synthetic Media and the EU AI Act’s Transparency Rules : the full Article 50 mechanics this guide’s closing section only summarizes.
- AI Transparency Obligations Across EU Regulations : how Article 50 fits alongside the Act’s other disclosure duties.
- Deepfake : the term’s precise legal and technical definition.
- AI Watermarking : how provenance marking actually works under the hood.
- Diffusion Models : the generative technique behind nearly every tool discussed here.
- Deepfake Fraud Reaches $3.7 Billion in Documented Losses : what unauthorized likeness generation looks like at scale, and why the consent question in this guide isn’t hypothetical.