Training Sand to Think: Artificial General Intelligence & Future of Physics

Training Sand to Think: Artificial General Intelligence & Future of Physics

The Evolution of Neural Networks and Large Language Models

Introduction to the Current Era

  • We are at a pivotal moment in civilization, having transformed sand into silicon chips and now using them to create neural networks capable of thinking.

Shift from Traditional Physics Papers

  • The speaker has authored around 40 theoretical physics papers but has shifted focus towards developing machines that can generate knowledge on an industrial scale. This change reflects a desire to contribute more significantly to scientific advancement.

Historical Context of Computational Tools

  • Previous computational tools like calculators were special-purpose, aiding specific tasks in physics. In contrast, large language models (LLMs) represent a significant leap as they can perform all aspects of a physicist's job.

Understanding Large Language Models

Definition and Capabilities

  • LLMs are general intelligences that can replace various roles within theoretical physics, marking a shift from traditional computing methods. They have capabilities beyond mere task-specific tools.

Recent Developments in LLM Technology

  • The speaker discusses the rapid advancements in LLM technology over the past five years, emphasizing their ability to perform complex mathematical and physical tasks effectively.

Characteristics of Large Language Models

Interaction with Users

  • Users can interact with LLMs through various platforms like Gemini and ChatGPT, which have surpassed the Turing test without much recognition for this achievement.

Structure and Growth of Neural Networks

  • Unlike traditional programs, neural networks grow by adjusting connections rather than being explicitly programmed; they start with random outputs that improve through training based on predictive accuracy.

Training Process for Neural Networks

Pre-training Phase

  • During pre-training, LLMs learn by predicting subsequent words based on vast amounts of text data; initial outputs may be nonsensical but improve significantly as training progresses through millions of words read.

Post-training Adjustments

  • After pre-training, models undergo post-training or "finishing school" where they are refined for politeness and helpfulness towards users rather than just word prediction accuracy. This stage is crucial for user interaction quality.

Scaling Laws in Machine Learning

Role of Physicists in AI Development

  • Physicists contributed significantly to understanding scaling laws which dictate how increasing model size or compute resources leads to better performance outcomes in machine learning systems. This empirical discovery initiated the modern boom in LLM development around 2020.

Implications of Scaling Laws

  • A linear relationship was found between increased compute resources spent on training neural networks and improved performance metrics; this insight attracted substantial investment into AI technologies from venture capitalists seeking returns on their investments based on these predictable outcomes.

Economic Factors Influencing AI Progress

Investment Trends

  • Exponential growth has been observed both in financial investment into AI training processes and computational power dedicated to these efforts since 2010, indicating strong market confidence in continued advancements within this field.

Algorithmic Innovations Driving Improvement

Importance of Algorithmic Progress

  • Beyond hardware improvements, algorithmic innovations have played a critical role in enhancing model efficiency; human ingenuity has led to significant reductions in inefficiencies during training processes over recent years.

Benchmarking Progress Over Time

Historical Performance Metrics

  • In 2019, large language models performed poorly compared to humans on benchmarks like MATH; however, rapid advancements have seen these models evolve dramatically within just four years.

MATH Benchmark Insights

  • Initial evaluations showed that even top-performing humans struggled with high school-level math problems while LLM performance was abysmal at only 6%. This highlighted significant challenges faced by early models regarding comprehension and parsing questions correctly.

Rapid Advancements Post-Benchmarking

  • Following initial struggles with benchmarks like MATH, newer systems quickly surpassed previous limitations achieving up to 90% accuracy shortly after introduction due largely to scaling effects combined with algorithmic improvements.

Future Directions for Large Language Models

Continuous Improvement Strategies

  • Techniques such as chain-of-thought prompting encourage deeper reasoning processes within models leading them toward higher accuracy levels when solving complex problems.

Collaborative Model Approaches

  • Engaging multiple large language models collaboratively enhances problem-solving capabilities further demonstrating potential pathways for future developments across diverse applications beyond mathematics alone.

Graduate Science and the GPQA Benchmark

Overview of GPQA

  • The GPQA benchmark simulates problems faced by first-year graduate students pursuing a PhD, focusing on mastery in their subject area.
  • PhD-level experts scored around 70% on this benchmark, indicating its complexity compared to high school mathematics.

Challenges with GPQA

  • Problems in GPQA require knowledge beyond adjacent fields; for example, a physicist may struggle with chemistry-related questions.
  • The performance of models improved significantly from random guessing to achieving near-perfect scores by early 2025, rendering the benchmark ineffective.

Skepticism About Learning vs. Memorization

Addressing Concerns

  • Critics argue that models might simply memorize answers found online rather than genuinely learning concepts.
  • To test this theory, researchers created look-alike problems not included in the original datasets to evaluate model performance on new challenges.

Evidence of Genuine Learning

  • Results showed little difference in performance between established test sets and newly invented problems, suggesting true understanding of math and physics concepts by large language models (LLMs).

Private Testing with Graduate Exams

Personal Benchmarking

  • The speaker created a private test set based on graduate exams from Stanford University to further assess LLM capabilities.
  • Over 18 months, these models achieved 100% accuracy on these exams, indicating significant advancements in their problem-solving abilities at the PhD level.

Evolution of Benchmarks

Popularity and Limitations

  • New benchmarks like "Humanity's Last Exam" emerged but quickly became obsolete as LLM performance surpassed expectations within a short time frame.
  • The International Maths Olympiad was initially thought too challenging for LLMs but was eventually passed successfully by them, demonstrating their advanced capabilities.

Creativity and Problem Solving

Achievements in Mathematics Competitions

  • LLM solutions for the International Maths Olympiad were noted for clarity and precision, resembling human-like reasoning processes rather than mere algorithmic outputs.

Classic Riddles and Model Responses

  • A classic riddle about gender assumptions illustrates how LLM responses can reflect societal biases while also showcasing cleverness when presented with variations of familiar problems.

Centaur-style Mathematical Research

Collaborative Efforts

  • A collaboration between humans and LLM led to novel mathematical research outputs that were recognized as significant contributions to the field.
  • Insights generated by LLM were deemed valuable enough for professional mathematicians to co-author papers alongside them.

Future Predictions for AI Development

Speculations on Progress

  • There are two potential paths: stagnation or continued advancement in AI capabilities.
  • Current limitations include low agency and slow learning rates; however, improvements have been observed over time.

Major Breakthrough Achieved

Milestone Achievement

  • Recently, an OpenAI model autonomously solved Erdős's unit distance conjecture—a major open problem—marking a significant milestone in AI mathematics.
  • This breakthrough is expected to pave the way for future successes as AI continues evolving beyond previous limitations.

Disanalogies Between Chess and Mathematics/Physics

Key Comparisons

  • The complexity of mathematics and physics surpasses that of chess, leading to a more extensive range of possibilities in these fields.
  • Computers excel at tactical execution, speed, and search capabilities but struggle with strategic thinking or "taste," a pattern observed in both chess and scientific endeavors.
  • Neural networks require significantly more games for training compared to humans; however, they can train much faster due to their ability to play numerous games quickly.

Training Dynamics of AI in Chess

Unique Features

  • Once trained, a chess bot does not need retraining like humans do; it continues improving beyond peak human performance without limitations.
  • Human players have improved due to learning from strong chess computers, resulting in today's best players being superior to historical champions despite still being weaker than the machines.

The Popularity of Chess and Future Implications

Potential for AI in Science

  • The rise of computer chess has coincided with increased popularity for the game itself; similar trends may emerge with large language models in scientific research.

Advancements in AI Intelligence

Cost Efficiency

  • Over recent years, not only has intelligence improved but the cost associated with producing fixed levels of intelligence has decreased dramatically.

Future Projections for AI Scientists

Implications for Physics

  • If one can create an "AI Einstein," scaling this technology could lead to billions of superhuman AIs contributing significantly to physics advancements.

Predictions for the Golden Era of Physics

Collaborative Renaissance

  • While predicting long-term outcomes is challenging due to rapid AI developments, the immediate future promises a collaborative renaissance between human experts and AI tools in science and mathematics.

Excitement Ahead for Physicists and Mathematicians

Anticipated Breakthroughs

  • This period is expected to be unprecedentedly exciting for physicists and mathematicians as many longstanding questions are anticipated to be resolved soon.
Video description

Our civilization has learned how to turn sand into silicon chips, silicon chips into neural networks, and neural networks into Artificial Intelligences (AIs). Over the last half-decade, the capabilities of large language model AIs (like ChatGPT and Gemini) have leapt from babbling preschoolers to International Math Olympiad gold medalists, and now beyond. This talk reviews recent progress in training AIs to do science and reasoning, and speculates as to what it will mean for the future of physics if these trends continue. About the Speaker Adam Brown leads Blueshift—a research team at Google DeepMind focused on advancing the scientific and reasoning capabilities of artificial intelligence—and is a core contributor to Gemini. Before Google, he studied physics and philosophy at Oxford, earned a PhD at Columbia, and subsequently held academic positions in the physics departments at Princeton and Stanford. There he taught Einstein’s general theory of relativity and conducted research on topics spanning the big bang, inflation, the multiverse, black holes, quantum computation, space elevators, bubbles of nothing, and the long-term fate of the universe, as well as the deep connections between physics and computer science. He joined Google in 2018. Stay in the loop: https://landing.perimeterinstitute.ca/newsletter-signup Support science: https://perimeterinstitute.ca/info/donors Follow Perimeter Institute: Facebook: https://www.facebook.com/pioutreach Instagram: https://www.instagram.com/perimeterinstitute LinkedIn: https://www.linkedin.com/company/perimeter-institute/ Bluesky: https://bsky.app/profile/perimeterinstitute.ca TikTok: https://www.tiktok.com/@perimeterinstitute