Claude Fable vs Opus 4.8: How Big is the Upgrade REALLY?

Claude Fable vs Opus 4.8: How Big is the Upgrade REALLY?

Comparing AI Models: Claude 3 Fable 5 vs. Opus

Introduction to the Challenge

  • The video introduces a head-to-head comparison between Anthropic's Claude 3 Fable 5 and another model, Opus, to build an app using identical prompts. The evaluation criteria include functionality, design, and cost.

Setting Ground Rules for Comparison

  • Three specific prompts will be used for both models:
  • A Product Requirement Document (PRD) detailing the app's features.
  • Design tweaks based on references provided by the presenter.
  • Adding additional functionalities to the created app.
  • No manual fixes are allowed during the process; if something breaks, it must self-correct. API key integration is permitted without counting as a prompt.

Conceptualizing the App Idea

  • The app aims to help entrepreneurs pitch their startup ideas effectively by providing feedback and humorous critiques similar to comedy roasts seen in popular shows. This dual approach combines entertainment with constructive criticism for better pitching skills.

Initial Setup of Testing Environments

Creating Test Folders

  • Two separate folders are created: one for Opus and another for Fable, ensuring that each model operates independently during testing sessions. The first prompt is input into both systems simultaneously.

Prompt Execution Begins

  • After setting up both environments, the presenter initiates the first prompt in each folder while preparing to monitor results closely from both models as they generate outputs concurrently.

First Results and Observations

Initial Outputs from Opus

  • Opus completes its task first, requesting an API key within its interface which allows users to utilize their own keys rather than relying on shared ones—this could reduce hosting costs significantly for developers.

Functionality Testing of Opus

  • The presenter tests Opus by pitching his company "We Are No Code," receiving a mix of humorous critiques and investor-focused feedback highlighting areas like market opportunity and business model clarity. Overall impressions indicate room for improvement but also some effective elements in functionality testing.

Evaluating Fable's Performance

Design Comparison with Fable

  • Upon testing Fable, initial observations reveal superior design aesthetics compared to Opus; it features smoother animations and a more polished user interface that enhances user experience significantly during interaction with the app functionalities.

Feedback from Pitching in Fable

  • Similar to Opus, when pitching "We Are No Code," feedback includes sharp humor alongside constructive insights about market positioning and differentiation strategies—indicating that while entertaining, it also provides valuable critique relevant for entrepreneurs seeking investment or customer engagement strategies.

Refining Designs Across Both Models

Design Improvement Prompts Issued

  • New prompts are issued focusing on refining designs further towards premium quality inspired by professional websites like Linear and Figma; this phase assesses how well each model can adapt based on detailed aesthetic instructions given previously by the presenter.

Final Feature Additions

Adding Investor Scorecard Functionality

  • A new feature request involves creating an investor scorecard assessing pitches based on criteria such as clarity and fundability; this step is crucial as deeper builds often lead to complications requiring robust functionality management within apps developed through these AI models.

Conclusion: Evaluating Overall Performance

Summary of Findings

  • In terms of functionality:
  • Fable received a score of 9/10 due to strong execution.
  • Opus scored lower at 7/10, indicating less effective performance overall despite some positive aspects noted.

Design Ratings:

  • Fable: Scored 7/10 reflecting good aesthetics but needing refinement.
  • Opus: Achieved an 8.5/10, praised for beautiful design elements.

Cost Analysis:

  • Token usage revealed significant differences:
  • Fable: Used approximately 21 million tokens.
  • Opus: Consumed around 42 million tokens, leading to higher operational costs under subscription plans.

The final verdict emphasizes that while both models have strengths in different areas (functionality vs design), there remains potential for improvement across all fronts before deployment into real-world applications is considered viable or cost-effective for non-coders looking at these tools as solutions in tech entrepreneurship contexts.[(1450)]

Video description

✅ FREE Startup Playbook: https://www.wearenocode.com/free-playbook 🎖️ FREE Community: https://members.wearenocode.com/join 🚀 Founder OS Course: https://www.wearenocode.com/founder-os 🛠 TOOLS AI Sales Bond: https://askbond.ai/ AI Coding: Lovable: https://dripl.ink/ZkTnR Hostinger: https://hostinger.com/wearenocode Emergent: https://app.emergent.sh/?via=wearenocode Base44: https://base44.pxf.io/c/4885934/2477538/25619?trafcat=hp Here are your three prompts for the Pitch Roast app: Prompt One: PRD "Build a web app called Pitch Roast that lets founders pitch their startup out loud and get AI feedback. The app should use the browser's built-in speech recognition (Web Speech API) to capture the user's voice live — include a prominent record button with a visual indicator showing it's actively listening, a live transcript that appears on screen as they speak, and a stop button to end the pitch. Once the pitch is captured, send the transcript to the Claude API with a prompt that generates two distinct outputs: first, a humorous roast of the pitch with a selectable intensity level (gentle, spicy, or brutal) chosen by the user before recording, and second, serious structured feedback written from the perspective of an institutional investor, covering market opportunity, business model clarity, differentiation, and the strength of the ask. Display the roast and investor feedback as two clearly separated sections. Include error handling for browsers that don't support speech recognition, API failures, and empty or too-short pitches. Add a loading state while the AI is generating feedback. Use semantic HTML, proper accessibility attributes, and make the layout fully responsive." Prompt Two: Design Polish "Refine the design to feel premium and entertaining, like a product from a modern startup studio. Take inspiration from Linear and Figma — clean whitespace, modern typography, subtle shadows, and smooth micro-interactions. The recording state should feel alive: add a pulsing animation on the record button and an animated waveform or sound-level indicator while the user is speaking. Style the roast section with playful visual personality (maybe a flame icon and warmer accent colors) and the investor feedback section with a serious, professional look (cooler tones, structured layout with clear subheadings). The intensity selector should be a fun, tactile toggle. Add fade-in animations when results load. The whole experience should feel polished enough that someone would share a screen recording of it on LinkedIn." Prompt Three: Scorecard Feature "Add an investor scorecard at the end of the feedback section that rates the pitch from 1 to 10 across four categories: clarity, market opportunity, differentiation, and fundability. Display the scores visually using animated progress bars or gauges, plus an overall verdict line at the top, something like 'Fundable with work' or 'Back to the drawing board.' The scores should be generated by the Claude API as part of the same analysis and returned as structured JSON. Also add a 'Pitch Again' button that resets the app so the user can immediately try an improved version of their pitch." ➡️ VISIT WeAreNoCode Website: https://www.wearenocode.com