Claude Fable 5 and GPT-5.6 Sol Music Video Challenge
- •Claude Fable 5 and GPT-5.6 Sol autonomously produced full music videos using specific tools and $25 or $100 budgets.
- •Claude Fable 5 resulted in higher total costs, reaching $73.65, while GPT-5.6 Sol maintained lower token-based operational expenses.
- •Models showed limitations in character consistency and story narrative, rarely utilizing self-review or iterative editing during production.
TryAI evaluated the autonomous creative capabilities of two advanced AI models, Claude Fable 5 and GPT-5.6 Sol, by tasking them to produce full music videos for "Uptown Funk" given a specific dollar budget. The experiment utilized an open-source agentic harness (a software framework for autonomous task execution) that provided models with six core tools: web search, budgeting, image generation, video generation, local shell commands via ffmpeg (a multimedia framework for processing video and audio), and a planning tool. Each model operated with complete autonomy to decide on research, generation techniques, and final editing.
Four distinct production runs were conducted, with budgets set at $25 and $100 per model. All runs successfully produced full-length, synchronized music videos. Results indicated that while Claude Fable 5 proved more expensive—totaling $73.65 for its $100 run—it completed tasks faster. GPT-5.6 Sol demonstrated lower token costs, remaining near $3 to $4 per run despite handling high token volumes. The $25 budget run for GPT-5.6 Sol uniquely utilized an image-to-video pipeline by generating stills first, whereas other runs relied primarily on text-to-video generation.
Performance analysis revealed significant limitations in current model reasoning for long-horizon creative tasks. All four models struggled with character consistency, tempo matching, and narrative coherence. While the agentic setup allowed models to automate the entire process, none demonstrated effective iterative refinement or self-review; they largely concatenated clips without returning to correct pacing issues or quality defects. Despite having access to both FAL and Replicate platforms, all models exclusively utilized FAL for generation. The full logs and open-source code are available at github.com/hershalb/music-video-arena for further inspection.