Falcon Soars to the Top - The NEW 40B LLM Rises above the rest.
Introduction to Falcon Language Model
In this section, the speaker introduces a new language model called Falcon. The model is trained from scratch with pre-training and has two versions: a 40 billion parameter model and a 7 billion parameter model.
Key Features of Falcon Language Model
- Falcon is not just a fine-tuning of another existing model but is trained from scratch with pre-training.
- There are two versions of the model: a 40 billion parameter model and a 7 billion parameter model.
- The 40 billion parameter version is one of the largest models available in the market.
- The model includes flash attention and other optimizations to speed up inference.
Benchmarking Results for Falcon Language Model
In this section, the speaker discusses benchmarking results for Falcon Language Model. The Hugging Face open large language model leaderboard shows that Falcon outperforms other models in various tasks.
Benchmarking Results for Falcon Language Model
- Hugging Face benchmarking systems show that Falcon outperforms other models in various tasks.
- For the arc reasoning challenge, HellaSwag, and other tasks, Falcon performs significantly better than other models.
- GPT 3.5 is currently performing slightly better than Falcon on few-shot tasks.
License for Using Falcon Language Model
In this section, the speaker discusses the license for using Falcon Language Model. If you make more than one million dollars per year using the model, you are expected to pay royalties.
License for Using Falcon Language Model
- Falcon Language Model is not free to use.
- If you make more than one million dollars per year using the model, you are expected to pay royalties.
Licensing and Model Release
In this section, the speaker discusses the licensing scheme for a 40 billion parameter-based model released by the Technology Innovation Institute. The royalty for using this model is 10%. The speaker also talks about other models that have been released.
Licensing Scheme
- The royalty for using the 40 billion parameter-based model is 10%.
- Some people are unhappy with this licensing scheme.
- The speaker thinks it's reasonable to charge a fee since training these models requires a lot of resources.
- The license terms are unclear, and it's uncertain if people will set up companies to provide access to the model as an API.
Model Release
- The Technology Innovation Institute has released several models, including a 40 billion parameter-based model and a Falcon 7 billion based model.
- These models were designed to outperform other language models on leaderboards.
- It's unclear why the Falcon 40B is recommended as a starting point for building instruct chat models.
Technical Details of Models
In this section, the speaker provides technical details about the Falcon 40B and Falcon 7B models released by the Technology Innovation Institute.
Falcon 40B Model
- This model has been fine-tuned on an unknown dataset.
- It has been submitted to leaderboards and performs well.
- The model is based on the GPT architecture and has 40 billion parameters.
- The speaker notes that the model's performance may not improve with further fine-tuning.
Falcon 7B Model
- This model was released on a mixture of chat instruct datasets.
- It's unclear if these datasets are distilled or can be used for training.
- The model uses the flash attention multi-query from Noam Shazeer's paper for the decoder.
Conclusion
In this section, the speaker concludes by discussing some interesting aspects of the models and expressing excitement about testing them out in code.
Final Thoughts
- The Technology Innovation Institute has released several language models, including a 40 billion parameter-based model and a Falcon 7 billion based model.
- It will be interesting to see how people use these models and whether they will improve with further fine-tuning.
- The speaker is excited to test out these models in code.
Setting up the Model
The speaker discusses setting up a custom colab to run the model and how they had to auto-cast everything to get it working.
Custom Colab Setup
- Running the model in a custom colab.
- The model has custom code for flash attention and decoding that is not compatible with eight-bit stuff.
- Auto-casting was necessary to get the model working.
Mixed Results
- Results are mixed, and the speaker is unsure about the quality of the dataset used for fine-tuning.
- Output quality is good for simple questions but disappointing for more complex ones.
Model Performance
The speaker discusses specific examples of how well or poorly the model performs on different types of questions.
Story Prompt Example
- The model writes a short story about a koala playing pool and beating all camelids. It's succinct but effective.
FLAN Reasoning Question Example
- The speaker is disappointed with one of the FLAN reasoning question responses, which should have been nine apples but was 11 instead.
- This makes them think that perhaps this dataset wasn't ideal for fine-tuning this particular model.
Haiku Example
- When asked to write a haiku in a single tweet, the response is disappointing as it can only provide formatting information.
- The speaker notes that this is not a good example of the model's capabilities.
Geoffrey Hinton Example
- The model understands that Geoffrey Hinton cannot have a conversation with George Washington due to the time difference but gets Hinton's birth year wrong.
Introduction
In this section, the speaker introduces the topic of discussion and highlights the strength of the model.
Interesting Findings
- The model shows strength in identifying that Harry Potter is a fictional character while Geoffrey Hinton is a real person.
- The model performs reasonably well on factual questions such as identifying Marcus Aurelius' son as Commodus.
Model Comparison
In this section, the speaker compares two models and discusses their performance on different tasks.
40 Billion vs. 7 Billion Model
- The 7 billion model writes a coherent argument when asked to write an email to Sam Altman, while the 40 billion model refuses to do so.
- The models may have been trained on different datasets, which could explain their varying performances.
- The reasoning question about apples stumps both models, with the 7 billion model providing an incorrect response.
Limitations of Fine-Tuning Models
In this section, the speaker discusses limitations of fine-tuning pre-trained models and suggests alternative approaches.
Limitations of Fine-Tuning Pre-Trained Models
- Fine-tuning pre-trained models may not always yield optimal results for specific tasks.
- Training custom models using Falcon 7B based models may be a better approach for certain tasks.
Future Directions
- It would be interesting to see how Falcon 7B based models perform when trained on better datasets for instruct fine-tuning.
Challenges of Serving the 40 Billion One
In this section, the speaker discusses the challenges of serving the 40 billion one model.
Challenges with 40 Billion One Model
- Serving the 40 billion one is going to be challenging.
- It's probably going to be beyond what most people want for this kind of model.
Fine Tuning with 7 Billion One Model
In this section, the speaker talks about how the 7 billion one model can be fine-tuned and provides some details about its training data.
Fine Tuning with 7 Billion One Model
- The 7 billion one model can certainly be played around with.
- The model has been trained on 1.5 trillion tokens, making it suitable for fine-tuning with filtered datasets used over the past month or so.
Conclusion and Call-to-Action
In this section, the speaker concludes by inviting questions from viewers and asking them to like and subscribe if they found the video useful.
Conclusion and Call-to-Action
- Viewers are invited to ask questions in comments below.
- Viewers are asked to like and subscribe if they found the video useful.
Turn any video into a summary like this
YouTube links, meetings, lectures — with transcripts, search, and chat.