In week 4 of the ENCORE build, I made a call that some advisors thought was wrong.
The call: do not build a custom video generation model. Route to existing vendors.
The reasoning has held up better than anything else in the architecture. This post is that reasoning.
The case for building a generator
The argument my advisors made was reasonable.
If you control the generation, you control the quality. You can tune for creator-specific use cases. You differentiate on output. You become a research-first company that ships better video than the labs.
It is the same argument every AI startup has made since 2021. Most of them are now wrappers, or acquired, or pivoting.
The case for the loop
My counter-argument was math.
Foundation labs (Google, OpenAI, ByteDance, Anthropic, Meta) are spending hundreds of millions per quarter on video generation models. They have access to data, compute, and talent that a startup cannot match.
The labs will keep improving. Veo 4 will be better than Veo 3. Sora 3 will be better than Sora 2. Kling 4 will be better than Kling 3. Each version arrives every 6 to 12 months.
If I build a custom generator, I am racing against this. Every release of mine has to outpace the labs' release cycle. The labs have 100x my budget. The math says I lose in 18 months.
If I route to all the labs, I get every improvement for free. I always have the best generator available. I never spend a dollar on model research.
The route-everywhere strategy is dominant if the moat lives somewhere else.
The full argument for why generation is commoditizing is here.
Where the moat actually lives
If the moat does not live in the generator, where does it live?
In the data layer. The loop. The Brand Brain. The accumulated per-user signal that no competitor can get by calling the same APIs I call.
The competitor with better generation has the same APIs I have. The competitor with my data has my data. The data is the moat. The generator is infrastructure.
This insight reframed the entire architecture. The product was no longer "ENCORE's AI editor." It was "ENCORE's data-driven content engine that uses the best AI generators in the world." Different product. Different team shape. Different funding story.
The team implications
Three concrete team decisions came from this call.
One. I did not need a head of ML research. I needed a head of data engineering. The hires reflected this. Our first engineering hire was a data infrastructure lead, not a generative AI researcher.
Two. Vendor relationships became core. I built direct relationships with Hedra, HeyGen, Seedance, Veo team, Sora team, Runway, Kling, Luma, ElevenLabs, and Suno. These relationships matter because we get early access to new models and pricing tiers.
Three. The product surface shifted. The interesting work was not "make the cut prettier." It was "make the orchestrator smarter at picking which vendor for which render." That is engineering, not research.
The fundraising implications
The decision changed how I raised.
If ENCORE were a video gen startup, I would be raising at $30M to $50M pre-money against a research thesis. Most investors at this stage are racing to back the next OpenAI of video. Crowded round. Limited room.
If ENCORE is a loop / data-layer startup that uses the best video gen as infrastructure, the raise is against a category-creating thesis. Different investors. Different milestones. Different long-term math.
The $3M seed at the loop thesis was easier to close than the same raise against a research thesis would have been. The investors who get it really get it. The ones who don't pass quickly. No middle.
The full operator playbook on this is in this post.
What we lost by not building a generator
Three real costs.
One. I cannot say "we built our own model." I have to explain the vendor router every time. The pitch is harder. The product is harder to demo for naive audiences. The press cycle is harder.
Two. Our gross margins are lower than a research-first competitor's. We pay vendors per render. They pay only for compute. Our margin on rendering is in the 35 to 40% range. A research-first company at scale could hit 70% to 80%.
Three. If a vendor jacks up prices or shuts down, we are exposed. We mitigate with multi-vendor routing (we can fail over) but the structural dependency is real.
These costs are real. They were the right trade.
What we gained
Four things that justify the trade.
One. Time to market. Shipped the v1 product in 11 weeks. A research-first build would have been 9 months minimum to get to comparable output.
Two. Best-of-breed quality. Every render uses the best vendor for that specific use case. A research-first company is stuck with their own model, even when it is not the best fit.
Three. Architecture that compounds. The moat lives in the data, which accumulates. The competitor with better generation cannot catch up on data without 90 days of users in their system.
Four. Strategic focus. I spend my time on the loop, not on the generator. The team's focus is on the data layer, the orchestrator, the analytics callbacks. We are not split between "build the model" and "build the product."
The category map of where the industry is converging is here.
What I would tell the next founder
If you are building in AI video right now, the question I would push you on is:
"What does your moat look like in 18 months when the foundation labs have closed every generation quality gap that exists today?"
If your answer is "we will still have better generation" you are betting against five companies with 100x your budget. I would not take that bet.
If your answer is "we will have accumulated per-user data that no competitor can replicate" you are betting on a different game. That is the game I picked.
5,200 creators are betting with me. Reserve your spot →