Case study · Symphony Clip Gen

Clip generation, rebuilt.

How a sprawl of AI generation tools became one creative platform — by refusing to design one more isolated app.

Role
Lead Designer
When
Jul – Oct 2025
Where
Symphony · TikTok
Focus
0→1 system design
Before — five separate apps, three histories, endless re-uploads; after — one card, one feed, zero downloads

Symphony is TikTok's generative-AI creative studio for advertisers. I led design for Clip Gen — the layer that turns a product photo, a prompt, or a reference clip into a finished video ad. I inherited it as a growing pile of separate mini-apps, one per capability. That fragmentation was the real design problem.

01 The sprawl

Adding motion took four steps.

Image-to-video, July 2025. Animating one still took four steps — and you lost every generation along the way.

Adding motion took four steps: pick type, pick image, confirm, then finally the tool

Users wanted something more dynamic.

A flat product shot was never the point — people wanted it in a scene, in motion, in a dozen variations. Every new ask bolted on another capability.

The flat product shot — the perfume on white
01 A product shot
Generated scene — perfume among tropical flowers
Generated scene — perfume on a sunset beach
Generated scene — perfume by the sea
02 Scenes it can generate
03 After i2v
04 What users wanted

Every capability became its own app.

Text-to-video, then image generation — each shipped as its own standalone app. Customers didn't work in a straight line; they wanted to generate, iterate, and loop back. The tool assumed a funnel; the work was a loop.

Input — a product photo of the perfume on a rock
01 Image-to-video

A young woman is unboxing a navy gift box. Natural lighting. Cozy interior.

02 Text-to-video
Product reference Scene reference Brand reference
Image-generation output — the composed Glaciera ad
03 Image generation
Product reference Gift box reference Model reference
04 Reference-to-video then coming

Underneath, the same six inputs.

Side by side, the three apps weren't different products — they were one input, configured six different ways.

01Model
02Upload rules
03Duration
04Generations
05Templates
06CTAs

Not a screen problem. An architecture problem.

The fix wasn't a cleaner layout. It was refusing to build the fourth tool — and designing the substrate every capability could sit on instead.

02 The system

Everything runs on one card.

I designed a single input card that every capability renders. What changes between them isn't the structure — it's configuration: the guidance, the controls, the default model. Capability became a setting, not an app.

The card configured for reference-to-video
Reference-to-video
The card configured for image-to-video
Image-to-video
The card configured for text-to-video
Text-to-video
The card configured for image generation
Image generation

Outputs become inputs.

Any output can feed straight back in as the next input — the loop customers had asked for, built for production volume, not a demo.

Image details — generate video straight from the output
From an image → generate video
Video details — send straight to Ads Manager
From a video → send to Ads Manager

One box, from one input to many.

The same card composes a couple of references or a dozen — new models slot into the same framework, so the surface never has to change.

The card with a couple of references
A few references
The same card with many references
Many references
The same card — a fully composed prompt
Fully composed

One photo in. A finished ad out.

One real run: a single product photo and one line of direction become a set of images, then a video — without ever leaving the card.

The input card with the jacket attached and the prompt: a woman wearing this in the streets of nyc
01 A photo and one line
Four generated images of a woman wearing the jacket on NYC streets
02 Four images
03 Then animated

Designing for uncertainty.

Generation takes time, and sometimes it fails. The wait had to feel handled — honest progress, a clear reason when something breaks, and an output you can act on the moment it lands.

Loading — 32% complete, check back in a few minutes Almost there — 95% Generation failed, with a clear reason and a retry A finished video, ready to use
03 Outcome

Clip Gen, at scale.

The unified card shipped across every capability. A year on, it's the engine behind most of Symphony's video output — and the next model is a config change, not a rebuild.

~70M

video generations a quarter — Q2 2026.

Generations by capability
R2V  36.3M I2V  28.3M T2V  5.1M
~687×
R2V growth at launch · Q1 → Q2 2026
~$15M
revenue per quarter

The next model just plugs in.

Because capability is configuration, the next one — dubbing, avatars, whatever's next — is a setting change, not a new app. The system was built to absorb models it hadn't met yet.

Image-to-videoJUL
+ Text-to-videoAUG
+ Image genOCT
+ Reference-to-videoNOV
One system2026