In 2022, you could build a billion-dollar company by being slightly better at AI video generation than the next team.
In 2026, you cannot.
The generators have commoditized. The wrappers around them have commoditized. Every founder who built on "we generate better" has either pivoted or is about to. The only defensive moat left in AI video is the loop. Not the generation step. The feedback loop that uses output to improve future output, per user, per brand, over time.
This post is the long-form argument for that thesis. With the proof points. With the implication for builders, investors, and anyone choosing tools in this space for the next 24 months.
The commoditization timeline
Three phases, roughly.
2022-2023: real differentiation. Runway, Pika, Stable Video, and a handful of others had meaningfully different capability surfaces. You could pick a generator and have a defensible quality advantage for 6 to 12 months.
2024: convergence begins. The foundation model labs started releasing video models. Veo, Sora, Kling, Hunyuan, Wan, Luma. Each one closed the gap to the prior best in a couple of months. Specialized startups began racing the labs and losing.
2025-2026: full commoditization. Every wrapper has API access to multiple foundation video models. The quality differences between Veo 3, Sora 2, Kling 3, and Seedance 2 are within rounding error for most use cases. The generators are now infrastructure, not differentiation. The wrappers are reselling them.
This is the same pattern photo apps went through between 2012 and 2016. By 2016, every photo app had the same filters because they were licensing the same underlying engines. The category collapsed into Instagram, which won because Instagram was a network, not a filter.
The editing-tool feature convergence I see today is the same pattern.
Why this matters for moats
A moat is a defensible position that competitors cannot replicate cheaply.
Generation is now cheap. Anyone can call the same APIs. The features built on top of those APIs (auto-caption, auto-cut, auto-music, auto-channel-variants) are also cheap because they are wrappers on the same underlying primitives.
If the input is commodity and the output is commodity, the moat has to live somewhere else. It has to live in something the competitor cannot get from an API call.
There are exactly three candidates:
-
Network effects. What I do makes the product better for you. Examples: TikTok, Reddit. Hard to build in a creator-tool context because creators do not naturally cluster around the tool.
-
Switching costs. What I do here makes it expensive for me to leave. Examples: Salesforce, Notion. Possible in creator tools if you anchor enough state per user.
-
Compounding intelligence per user. What I do here makes the system smarter about me specifically. Examples: Spotify's recommendation engine, Netflix's. Possible in creator tools if you have the right data layer.
ENCORE's bet is on (2) and (3) combined, with (1) playing a secondary role.
The full breakdown of why we picked the loop over a better generator is in this post.
What the loop actually is
The loop is the feedback structure that connects four phases per user:
Creation → Publishing → Analytics → Better Creation.
Phase 1: the user creates content via the tool.
Phase 2: the tool publishes that content across channels.
Phase 3: the tool ingests channel-specific performance analytics (views, retention, CTR, engagement, audience demo).
Phase 4: the tool uses those analytics to shape the next creation prompt for that specific user, that specific brand, that specific audience.
The loop is closed. Output feeds back into input. Every iteration makes the next one stronger.
This is what I mean when I say the loop is the moat. The loop accumulates value per user per week of use. After 8 weeks, the loop's accumulated signal for a specific creator is something no competitor can replicate without 8 weeks of that creator's data. Even if the competitor has better generation. Even if the competitor has cheaper pricing. Even if the competitor has a slicker UI.
The data is the moat. The loop is what makes the data accumulate.
A more detailed primer on the Brand Brain that runs the loop is here.
Why generation alone cannot do this
Generation is a one-way function. It takes a prompt, returns an output, has no memory.
You can stack generation calls together. You can chain them. You can RAG them with reference videos. You can do all of this and still not have a loop, because the loop requires four things generation does not provide:
-
Per-user state that persists across renders. Generation does not have this. Each call is stateless.
-
Multi-channel publishing infrastructure. Generation does not have this. It needs a separate publishing layer.
-
Analytics ingestion from the platforms. Generation does not have this. It needs API access to YouTube, TikTok, Instagram, LinkedIn, Facebook, Pinterest performance data.
-
A scoring and learning mechanism that updates the per-user state based on analytics. Generation does not have this. It needs a recommendation engine.
To build a loop, you need all four. Generation is one quarter of the moat at best.
This is why every "AI video generation" startup is racing toward a wall. They are perfecting the one quarter. The other three quarters require completely different infrastructure that they have not built.
The category map of where generation, scheduling, and analytics are converging is here.
The competitor landscape, mapped
Three buckets of competition today.
Bucket 1: pure generators. Pika, Runway, Sora, Veo, Kling, Seedance, Luma. They are racing each other on output quality. None of them are building the loop because that is not their business. They are infrastructure for everyone, including ENCORE. We use them all.
Bucket 2: wrappers + features. CapCut, Submagic, OpusClip, Veed, Captions, Descript, HeyGen. They wrap one or more generators and add features. They are converging on the same feature set and racing each other on price. None of them are building the loop because they were built around a single-phase wrapper architecture, not a multi-phase loop.
Bucket 3: schedulers + analytics. Later, Buffer, Hootsuite, Sprout, Metricool, Iconosquare. They publish and analyze but do not generate. They cannot close the loop because the loop requires a unified data layer across creation and analytics.
Each bucket has 10 to 30 competitors. None of them have all three.
ENCORE is the loop. We do the generation (via vendor APIs we route to), the wrapping (per-render brand-kit application + channel variants), and the publish + analytics in one system, with one shared data layer.
What this means for builders
If you are building in AI video right now, three implications.
One. Do not build a pure generator. The foundation labs will catch you in 12 months.
Two. Do not build a pure wrapper. The wrapper category is already converging, with prices falling. You are racing into a commoditized market with no defensible state.
Three. Build something that compounds. If you cannot articulate why your product is better for your user on day 90 than it is on day 1, you do not have a moat. The bar for "better on day 90" is whether the system has accumulated user-specific signal that makes the next interaction smarter than the last.
What this means for creators and agencies
If you are picking tools, three implications.
One. The current 12-tool stack you are paying for is racing toward a wall. Each tool is becoming a feature in something bigger. Your switching cost away from the stack is dropping every quarter as the system that absorbs it gets better.
Two. The single move that compounds is to pick the system with the strongest data layer. Not the slickest UI. Not the cheapest price. The one where your second month is meaningfully better than your first month because the system learned about you.
Three. You have a 6 to 12 month window to anchor your Brand Brain inside a loop system. After that, the leaders will have accumulated enough per-user data that switching gets meaningfully expensive. The window matters.
What this means for investors
If you are looking at AI video deals, the question is not "how good is the generation."
The question is "what does the system know about a user on day 90 that no competitor knows."
If the answer is "the same thing every competitor knows" you are looking at a wrapper. Pass.
If the answer is "per-user voice patterns, color preferences, hook formulas, pacing rhythms, caption styles, and audience signal vectors that have been refined over 90 days of use" you are looking at a loop. That is the moat.
What I would tell the room
Generation will continue to improve. Veo 4 will be better than Veo 3. Sora 3 will be better than Sora 2. The foundation models will keep getting cheaper and faster.
None of that changes the moat thesis. The generation layer is infrastructure. Infrastructure is not a business in the long run. It is a cost center.
The business is the loop. The product is the loop. The category is the loop.
5,200 creators and 12 agencies on the waitlist agree.
If you want the founder version of this argument, the operator notes from 11 weeks of building ENCORE are in this post.