I Tested the Fable 5 Killer (Hermes Agent)
Exploring New AI Models: Sakana Fugu and GLM 5.2
Introduction to the Models
- In the past week, two models, Sakana Fugu and GLM 5.2, have emerged as contenders against Fable 5.
- The rapid evolution of AI models can be overwhelming; hence, testing these models is essential to determine their effectiveness.
- Sakana Fugu acts as an intelligent router for various Frontier models, while GLM 5.2 is a cost-effective open-weight model from China.
Performance Insights
- Users report that GLM 5.2 offers better performance than Opus at a significantly lower cost.
- Sakana Fugu intelligently routes questions to the most suitable model based on task type, enhancing efficiency in responses.
Understanding Sakana Fugu
- It operates through one API that manages multiple underlying models, allowing dynamic selection based on user queries.
- Described as a multi-agent system delivered as one model, it aims to compete directly with Fable 5.
Overview of GLM 5.2
- Known for its affordability (one-sixth the price of competitors), GLM 5.2 boasts a million context window and can be run locally or within Hermes agents.
- Its performance metrics are crucial for evaluating cost-effectiveness alongside output quality.
Connecting Models to Hermes Agent
Setting Up Sakana Fugu
- Instructions are provided for connecting Sakana to Hermes agent via API keys obtained from the Sakana website.
- Users can securely input API keys into Hermes by generating terminal commands through conversational prompts.
Testing Model Capabilities
- Initial tests involve querying both models about simple tasks like counting letters in words and providing creative descriptions.
Dynamic Model Selection in Hermes
Advantages of Hermes Agent
- The ability to dynamically switch between different AI models using voice commands enhances usability and flexibility during interactions.
Setting Up GLM 5.2
Connection Process
- A guide is available for obtaining a direct API key for GLM 5.2 or connecting it via Open Router for ease of use.
Comparative Testing Between Models
Evaluating Responses
- Queries posed include assessing sports preferences and basic arithmetic tasks across all three tested models (GLM 5.2, Opus 4.8, and Sukuna).
Tool Calling Test Results
Performance Analysis
- A test involving retrieving email subjects reveals varying success rates among the three models; Sukuna performs best initially but requires retries from GLM to succeed.
Website Creation Challenge
Creative Output Assessment
- Each model is tasked with creating a visually appealing one-page website for sparkling water; results show varied levels of creativity and execution quality.
Final Evaluation of Model Efficiency
Latency and Token Usage
- Sakana: Wins initial task but has high token usage compared to others.
- GLM: Fastest response time despite some failures in tool calling; excels in coding tasks.
- Opus: Performs adequately but not impressively compared to others regarding speed and token efficiency.
Conclusion on Model Effectiveness
- Overall findings suggest that while none surpasses Claude's capabilities entirely, GLM shows significant promise due to its low cost-to-performance ratio ().