Apple’s Hidden AI Model… The Speed Test Apple Never Showed
Exploring the Hidden AI in Mac OS
Introduction to the New AI Tool
- A new AI feature is integrated into Mac OS, specifically in the terminal, which many users are unaware of.
- The speaker tests this tool on a Mac Mini upgraded to Mac OS 27 beta and compares its performance with an M3 Ultra Mac Studio.
Overview of Super Nori
- Introduction to Super Nori, a proactive family AI agent designed to operate in the background within family contexts.
- Unlike traditional tools that require user prompts, Super Nori identifies issues autonomously and suggests solutions proactively.
- Examples include booking rides during traffic jams or making dinner reservations when calendar changes threaten plans.
Understanding FM Tool in Mac OS 27
- The FM tool is a built-in command-line interface (CLI) that utilizes large language models (LLMs).
- Users can access it directly from their terminal without needing additional installations like Olama or LM Studio.
Benchmarking Performance of FM Tool
Initial Testing Challenges
- Initial benchmarks using Llama Beni indicated unexpected processing speeds for Apple's model, suggesting potential errors in measurement.
- Apple's unique server setup complicates standard benchmarking as it does not fully comply with OpenAI standards.
Development of Custom Benchmarking Tool
- Due to limitations with existing tools, the speaker created a custom benchmark called Apple FM bench for accurate measurements.
- Results showed local processing speeds of approximately 50 tokens per second for prompt processing and 52 for decoding.
Comparing Local vs. Cloud Performance
Performance Insights
- The cloud version of FM operates about three times faster than the local version, but local performance remains crucial for hardware efficiency.
Hardware Comparison: Mac Mini vs. M3 Ultra
- The speaker compares memory bandwidth between the M4 Pro (Mac Mini at 273 GB/s) and M3 Ultra (819 GB/s), hypothesizing better performance from more powerful hardware.
Unexpected Benchmark Results
Analyzing Processing Speeds
- Surprisingly similar results were observed between both machines despite significant hardware differences; prompting further investigation into underlying causes.
Investigating Neural Engine Usage
- Speculation arose that both systems might be utilizing the same neural engine architecture due to identical performance metrics across different models.
Testing Neural Engine Functionality
Measurement Limitations Encountered
- Attempts to monitor neural engine usage during testing revealed zero activity, leading to doubts about its involvement in processing tasks.
Confirming Neural Engine's Role
- Tests confirmed that enabling CoreML’s neural engine significantly improved processing speed by over two times compared to running without it.
Final Conclusions on Performance Dynamics
Determining Resource Allocation
- By stressing both the neural engine and GPU while running FM, it was determined that decoding relies on the neural engine while prompt processing uses GPU resources effectively.
Cost Efficiency Insights
- Ultimately concluded that using Apple's built-in FM tool provides equivalent performance on a lower-cost Mac Mini compared to high-end models like the M3 Ultra Max Studio.
Future Considerations
Potential Developments
- Noted that since this feature is still in beta, future updates could alter its functionality or resource allocation strategies.
- Encouraged viewers to explore their own experiences with Apple’s hidden features and provided links for further exploration and community feedback.